<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Bayesian network</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Bayesian_network"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/ext.math.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Bayesian_network rootpage-Bayesian_network skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Bayesian network</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<style data-mw-deduplicate="TemplateStyles:r1251242444">
/* start https://en.wikipedia.org/ */
.mw-parser-output .ambox{border:1px solid #a2a9b1;border-left:10px solid #36c;background-color:#fbfbfb;box-sizing:border-box}.mw-parser-output .ambox+link+.ambox,.mw-parser-output .ambox+link+style+.ambox,.mw-parser-output .ambox+link+link+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+style+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+link+.ambox{margin-top:-1px}html body.mediawiki .mw-parser-output .ambox.mbox-small-left{margin:4px 1em 4px 0;overflow:hidden;width:238px;border-collapse:collapse;font-size:88%;line-height:1.25em}.mw-parser-output .ambox-speedy{border-left:10px solid #b32424;background-color:#fee7e6}.mw-parser-output .ambox-delete{border-left:10px solid #b32424}.mw-parser-output .ambox-content{border-left:10px solid #f28500}.mw-parser-output .ambox-style{border-left:10px solid #fc3}.mw-parser-output .ambox-move{border-left:10px solid #9932cc}.mw-parser-output .ambox-protection{border-left:10px solid #a2a9b1}.mw-parser-output .ambox .mbox-text{border:none;padding:0.25em 0.5em;width:100%}.mw-parser-output .ambox .mbox-image{border:none;padding:2px 0 2px 0.5em;text-align:center}.mw-parser-output .ambox .mbox-imageright{border:none;padding:2px 0.5em 2px 0;text-align:center}.mw-parser-output .ambox .mbox-empty-cell{border:none;padding:0;width:1px}.mw-parser-output .ambox .mbox-image-div{width:52px}@media(min-width:720px){.mw-parser-output .ambox{margin:0 10%}}@media print{body.ns-0 .mw-parser-output .ambox{display:none!important}}
/* end https://en.wikipedia.org/ */
</style>
<style data-mw-deduplicate="TemplateStyles:r1129693374">
/* start https://en.wikipedia.org/ */
.mw-parser-output .hlist dl,.mw-parser-output .hlist ol,.mw-parser-output .hlist ul{margin:0;padding:0}.mw-parser-output .hlist dd,.mw-parser-output .hlist dt,.mw-parser-output .hlist li{margin:0;display:inline}.mw-parser-output .hlist.inline,.mw-parser-output .hlist.inline dl,.mw-parser-output .hlist.inline ol,.mw-parser-output .hlist.inline ul,.mw-parser-output .hlist dl dl,.mw-parser-output .hlist dl ol,.mw-parser-output .hlist dl ul,.mw-parser-output .hlist ol dl,.mw-parser-output .hlist ol ol,.mw-parser-output .hlist ol ul,.mw-parser-output .hlist ul dl,.mw-parser-output .hlist ul ol,.mw-parser-output .hlist ul ul{display:inline}.mw-parser-output .hlist .mw-empty-li{display:none}.mw-parser-output .hlist dt::after{content:": "}.mw-parser-output .hlist dd::after,.mw-parser-output .hlist li::after{content:" · ";font-weight:bold}.mw-parser-output .hlist dd:last-child::after,.mw-parser-output .hlist dt:last-child::after,.mw-parser-output .hlist li:last-child::after{content:none}.mw-parser-output .hlist dd dd:first-child::before,.mw-parser-output .hlist dd dt:first-child::before,.mw-parser-output .hlist dd li:first-child::before,.mw-parser-output .hlist dt dd:first-child::before,.mw-parser-output .hlist dt dt:first-child::before,.mw-parser-output .hlist dt li:first-child::before,.mw-parser-output .hlist li dd:first-child::before,.mw-parser-output .hlist li dt:first-child::before,.mw-parser-output .hlist li li:first-child::before{content:" (";font-weight:normal}.mw-parser-output .hlist dd dd:last-child::after,.mw-parser-output .hlist dd dt:last-child::after,.mw-parser-output .hlist dd li:last-child::after,.mw-parser-output .hlist dt dd:last-child::after,.mw-parser-output .hlist dt dt:last-child::after,.mw-parser-output .hlist dt li:last-child::after,.mw-parser-output .hlist li dd:last-child::after,.mw-parser-output .hlist li dt:last-child::after,.mw-parser-output .hlist li li:last-child::after{content:")";font-weight:normal}.mw-parser-output .hlist ol{counter-reset:listitem}.mw-parser-output .hlist ol>li{counter-increment:listitem}.mw-parser-output .hlist ol>li::before{content:" "counter(listitem)"\a0 "}.mw-parser-output .hlist dd ol>li:first-child::before,.mw-parser-output .hlist dt ol>li:first-child::before,.mw-parser-output .hlist li ol>li:first-child::before{content:" ("counter(listitem)"\a0 "}
/* end https://en.wikipedia.org/ */
</style><style data-mw-deduplicate="TemplateStyles:r1246091330">
/* start https://en.wikipedia.org/ */
.mw-parser-output .sidebar{width:22em;float:right;clear:right;margin:0.5em 0 1em 1em;background:var(--background-color-neutral-subtle,#f8f9fa);border:1px solid var(--border-color-base,#a2a9b1);padding:0.2em;text-align:center;line-height:1.4em;font-size:88%;border-collapse:collapse;display:table}body.skin-minerva .mw-parser-output .sidebar{display:table!important;float:right!important;margin:0.5em 0 1em 1em!important}.mw-parser-output .sidebar-subgroup{width:100%;margin:0;border-spacing:0}.mw-parser-output .sidebar-left{float:left;clear:left;margin:0.5em 1em 1em 0}.mw-parser-output .sidebar-none{float:none;clear:both;margin:0.5em 1em 1em 0}.mw-parser-output .sidebar-outer-title{padding:0 0.4em 0.2em;font-size:125%;line-height:1.2em;font-weight:bold}.mw-parser-output .sidebar-top-image{padding:0.4em}.mw-parser-output .sidebar-top-caption,.mw-parser-output .sidebar-pretitle-with-top-image,.mw-parser-output .sidebar-caption{padding:0.2em 0.4em 0;line-height:1.2em}.mw-parser-output .sidebar-pretitle{padding:0.4em 0.4em 0;line-height:1.2em}.mw-parser-output .sidebar-title,.mw-parser-output .sidebar-title-with-pretitle{padding:0.2em 0.8em;font-size:145%;line-height:1.2em}.mw-parser-output .sidebar-title-with-pretitle{padding:0.1em 0.4em}.mw-parser-output .sidebar-image{padding:0.2em 0.4em 0.4em}.mw-parser-output .sidebar-heading{padding:0.1em 0.4em}.mw-parser-output .sidebar-content{padding:0 0.5em 0.4em}.mw-parser-output .sidebar-content-with-subgroup{padding:0.1em 0.4em 0.2em}.mw-parser-output .sidebar-above,.mw-parser-output .sidebar-below{padding:0.3em 0.8em;font-weight:bold}.mw-parser-output .sidebar-collapse .sidebar-above,.mw-parser-output .sidebar-collapse .sidebar-below{border-top:1px solid #aaa;border-bottom:1px solid #aaa}.mw-parser-output .sidebar-navbar{text-align:right;font-size:115%;padding:0 0.4em 0.4em}.mw-parser-output .sidebar-list-title{padding:0 0.4em;text-align:left;font-weight:bold;line-height:1.6em;font-size:105%}.mw-parser-output .sidebar-list-title-c{padding:0 0.4em;text-align:center;margin:0 3.3em}@media(max-width:640px){body.mediawiki .mw-parser-output .sidebar{width:100%!important;clear:both;float:none!important;margin-left:0!important;margin-right:0!important}}body.skin--responsive .mw-parser-output .sidebar a>img{max-width:none!important}@media screen{html.skin-theme-clientpref-night .mw-parser-output .sidebar:not(.notheme) .sidebar-list-title,html.skin-theme-clientpref-night .mw-parser-output .sidebar:not(.notheme) .sidebar-title-with-pretitle{background:transparent!important}html.skin-theme-clientpref-night .mw-parser-output .sidebar:not(.notheme) .sidebar-title-with-pretitle a{color:var(--color-progressive)!important}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .sidebar:not(.notheme) .sidebar-list-title,html.skin-theme-clientpref-os .mw-parser-output .sidebar:not(.notheme) .sidebar-title-with-pretitle{background:transparent!important}html.skin-theme-clientpref-os .mw-parser-output .sidebar:not(.notheme) .sidebar-title-with-pretitle a{color:var(--color-progressive)!important}}@media print{body.ns-0 .mw-parser-output .sidebar{display:none!important}}
/* end https://en.wikipedia.org/ */
</style><table class="sidebar nomobile nowraplinks hlist"><tbody><tr><td class="sidebar-pretitle">Part of a series on</td></tr><tr><th class="sidebar-title-with-pretitle"><a href="Bayesian_statistics" title="Bayesian statistics">Bayesian statistics</a></th></tr><tr><td class="sidebar-image"><span typeof="mw:File"></span></td></tr><tr><td class="sidebar-content">
<a href="Posterior_probability" title="Posterior probability">Posterior</a> = <a href="Likelihood_function" title="Likelihood function">Likelihood</a> × <a href="Prior_probability" title="Prior probability">Prior</a> ÷ <a href="Marginal_likelihood" title="Marginal likelihood">Evidence</a></td>
</tr><tr><th class="sidebar-heading">
Background</th></tr><tr><td class="sidebar-content">
<ul><li><a href="Bayesian_inference" title="Bayesian inference">Bayesian inference</a></li>
<li><a href="Bayesian_probability" title="Bayesian probability">Bayesian probability</a></li>
<li><a href="Bayes'_theorem" title="Bayes' theorem">Bayes' theorem</a></li>
<li><a href="Bernstein%E2%80%93von_Mises_theorem" title="Bernstein–von Mises theorem">Bernstein–von Mises theorem</a></li>
<li><a href="Coherence_(philosophical_gambling_strategy)" class="mw-redirect" title="Coherence (philosophical gambling strategy)">Coherence</a></li>
<li><a href="Cox's_theorem" title="Cox's theorem">Cox's theorem</a></li>
<li><a href="Cromwell's_rule" title="Cromwell's rule">Cromwell's rule</a></li>
<li><a href="Likelihood_principle" title="Likelihood principle">Likelihood principle</a></li>
<li><a href="Principle_of_indifference" title="Principle of indifference">Principle of indifference</a></li>
<li><a href="Principle_of_maximum_entropy" title="Principle of maximum entropy">Principle of maximum entropy</a></li></ul></td>
</tr><tr><th class="sidebar-heading">
Model building</th></tr><tr><td class="sidebar-content">
<ul><li><a href="Conjugate_prior" title="Conjugate prior">Conjugate prior</a></li>
<li><a href="Bayesian_linear_regression" title="Bayesian linear regression">Linear regression</a></li>
<li><a href="Empirical_Bayes_method" title="Empirical Bayes method">Empirical Bayes</a></li>
<li><a href="Bayesian_hierarchical_modeling" title="Bayesian hierarchical modeling">Hierarchical model</a></li></ul></td>
</tr><tr><th class="sidebar-heading">
Posterior approximation</th></tr><tr><td class="sidebar-content">
<ul><li><a href="Markov_chain_Monte_Carlo" title="Markov chain Monte Carlo">Markov chain Monte Carlo</a></li>
<li><a href="Laplace's_approximation" title="Laplace's approximation">Laplace's approximation</a></li>
<li><a href="Integrated_nested_Laplace_approximations" title="Integrated nested Laplace approximations">Integrated nested Laplace approximations</a></li>
<li><a href="Variational_Bayesian_methods" title="Variational Bayesian methods">Variational inference</a></li>
<li><a href="Approximate_Bayesian_computation" title="Approximate Bayesian computation">Approximate Bayesian computation</a></li></ul></td>
</tr><tr><th class="sidebar-heading">
Estimators</th></tr><tr><td class="sidebar-content">
<ul><li><a href="Bayesian_estimator" class="mw-redirect" title="Bayesian estimator">Bayesian estimator</a></li>
<li><a href="Credible_interval" title="Credible interval">Credible interval</a></li>
<li><a href="Maximum_a_posteriori_estimation" title="Maximum a posteriori estimation">Maximum a posteriori estimation</a></li></ul></td>
</tr><tr><th class="sidebar-heading">
Evidence approximation</th></tr><tr><td class="sidebar-content">
<ul><li><a href="Evidence_lower_bound" title="Evidence lower bound">Evidence lower bound</a></li>
<li><a href="Nested_sampling_algorithm" title="Nested sampling algorithm">Nested sampling</a></li></ul></td>
</tr><tr><th class="sidebar-heading">
Model evaluation</th></tr><tr><td class="sidebar-content">
<ul><li><a href="Bayes_factor" title="Bayes factor">Bayes factor</a> (<a href="Bayesian_information_criterion" title="Bayesian information criterion">Schwarz criterion</a>)</li>
<li><a href="Bayesian_model_averaging" class="mw-redirect" title="Bayesian model averaging">Model averaging</a></li>
<li><a href="Posterior_predictive_distribution" title="Posterior predictive distribution">Posterior predictive</a></li></ul></td>
</tr><tr><td class="sidebar-below">
<ul><li><span class="nowrap"><span class="skin-invert-image noviewer" typeof="mw:File"></span> </span><a href="Portal%3AMathematics" title="Portal:Mathematics">Mathematics portal</a></li></ul></td></tr><tr><td class="sidebar-navbar"><style data-mw-deduplicate="TemplateStyles:r1239400231">
/* start https://en.wikipedia.org/ */
.mw-parser-output .navbar{display:inline;font-size:88%;font-weight:normal}.mw-parser-output .navbar-collapse{float:left;text-align:left}.mw-parser-output .navbar-boxtext{word-spacing:0}.mw-parser-output .navbar ul{display:inline-block;white-space:nowrap;line-height:inherit}.mw-parser-output .navbar-brackets::before{margin-right:-0.125em;content:"[ "}.mw-parser-output .navbar-brackets::after{margin-left:-0.125em;content:" ]"}.mw-parser-output .navbar li{word-spacing:-0.125em}.mw-parser-output .navbar a>span,.mw-parser-output .navbar a>abbr{text-decoration:inherit}.mw-parser-output .navbar-mini abbr{font-variant:small-caps;border-bottom:none;text-decoration:none;cursor:inherit}.mw-parser-output .navbar-ct-full{font-size:114%;margin:0 7em}.mw-parser-output .navbar-ct-mini{font-size:114%;margin:0 4em}html.skin-theme-clientpref-night .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}@media(prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}}@media print{.mw-parser-output .navbar{display:none!important}}
/* end https://en.wikipedia.org/ */
</style></td></tr></tbody></table>
<p>A <b>Bayesian network</b> (also known as a <b>Bayes network</b>, <b>Bayes net</b>, <b>belief network</b>, or <b>decision network</b>) is a <a href="Probabilistic_graphical_model" class="mw-redirect" title="Probabilistic graphical model">probabilistic graphical model</a> that represents a set of variables and their <a href="Conditional_dependence" title="Conditional dependence">conditional dependencies</a> via a <a href="Directed_acyclic_graph" title="Directed acyclic graph">directed acyclic graph</a> (DAG).<sup id="cite_ref-1" class="reference"><a href="#cite_note-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> While it is one of several forms of <a href="Causal_notation" title="Causal notation">causal notation</a>, causal networks are special cases of Bayesian networks. Bayesian networks are ideal for taking an event that occurred and predicting the likelihood that any one of several possible known causes was the contributing factor. For example, a Bayesian network could represent the probabilistic relationships between diseases and symptoms. Given symptoms, the network can be used to compute the probabilities of the presence of various diseases.
</p><p>Efficient algorithms can perform <a href="Inference" title="Inference">inference</a> and <a href="Machine_learning" title="Machine learning">learning</a> in Bayesian networks. Bayesian networks that model sequences of variables (<i>e.g.</i> <a href="Speech_recognition" title="Speech recognition">speech signals</a> or <a href="Peptide_sequence" class="mw-redirect" title="Peptide sequence">protein sequences</a>) are called <a href="Dynamic_Bayesian_network" title="Dynamic Bayesian network">dynamic Bayesian networks</a>. Generalizations of Bayesian networks that can represent and solve decision problems under uncertainty are called <a href="Influence_diagram" title="Influence diagram">influence diagrams</a>.
</p>
<style data-mw-deduplicate="TemplateStyles:r886046785">
/* start https://en.wikipedia.org/ */
.mw-parser-output .toclimit-2 .toclevel-1 ul,.mw-parser-output .toclimit-3 .toclevel-2 ul,.mw-parser-output .toclimit-4 .toclevel-3 ul,.mw-parser-output .toclimit-5 .toclevel-4 ul,.mw-parser-output .toclimit-6 .toclevel-5 ul,.mw-parser-output .toclimit-7 .toclevel-6 ul{display:none}
/* end https://en.wikipedia.org/ */
</style><div class="toclimit-3"><meta property="mw:PageProp/toc"></div>
<div class="mw-heading mw-heading2"><h2 id="Graphical_model">Graphical model</h2></div>
<p>Formally, Bayesian networks are <a href="Directed_acyclic_graph" title="Directed acyclic graph">directed acyclic graphs</a> (DAGs) whose nodes represent variables in the <a href="Bayesian_probability" title="Bayesian probability">Bayesian</a> sense: they may be observable quantities, <a href="Latent_variable" class="mw-redirect" title="Latent variable">latent variables</a>, unknown parameters or hypotheses. Each edge represents a direct conditional dependency. Any pair of nodes that are not connected (i.e. no path connects one node to the other) represent variables that are <a href="Conditional_independence" title="Conditional independence">conditionally independent</a> of each other. Each node is associated with a <a href="Probability_distribution" title="Probability distribution">probability function</a> that takes, as input, a particular set of values for the node's <a href="Glossary_of_graph_theory#Directed_acyclic_graphs" title="Glossary of graph theory">parent</a> variables, and gives (as output) the probability (or probability distribution, if applicable) of the variable represented by the node. For example, if <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle m}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>m</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle m}</annotation>
</semantics>
</math></span><img src="./0a07d98bb302f3856cbabc47b2b9016692e3f7bc.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:2.04ex; height:1.676ex;" alt="{\displaystyle m}" loading="lazy"></span> parent nodes represent <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle m}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>m</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle m}</annotation>
</semantics>
</math></span><img src="./0a07d98bb302f3856cbabc47b2b9016692e3f7bc.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:2.04ex; height:1.676ex;" alt="{\displaystyle m}" loading="lazy"></span> <a href="Boolean_data_type" title="Boolean data type">Boolean variables</a>, then the probability function could be represented by a table of <small><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle 2^{m}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msup>
<mn>2</mn>
<mrow class="MJX-TeXAtom-ORD">
<mi>m</mi>
</mrow>
</msup>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle 2^{m}}</annotation>
</semantics>
</math></span><img src="./667d0154f26e56e3f7979803f08afac16b4dcb16.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:2.837ex; height:2.343ex;" alt="{\displaystyle 2^{m}}" loading="lazy"></span></small> entries, one entry for each of the <small><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle 2^{m}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msup>
<mn>2</mn>
<mrow class="MJX-TeXAtom-ORD">
<mi>m</mi>
</mrow>
</msup>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle 2^{m}}</annotation>
</semantics>
</math></span><img src="./667d0154f26e56e3f7979803f08afac16b4dcb16.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:2.837ex; height:2.343ex;" alt="{\displaystyle 2^{m}}" loading="lazy"></span></small> possible parent combinations. Similar ideas may be applied to undirected, and possibly cyclic, graphs such as <a href="Markov_network" class="mw-redirect" title="Markov network">Markov networks</a>.
</p>
<div class="mw-heading mw-heading2"><h2 id="Example">Example</h2></div>
<p>Suppose we want to model the dependencies between three variables: the sprinkler (or more appropriately, its state - whether it is on or not), the presence or absence of rain and whether the grass is wet or not. Observe that two events can cause the grass to become wet: an active sprinkler or rain. Rain has a direct effect on the use of the sprinkler (namely that when it rains, the sprinkler usually is not active). This situation can be modeled with a Bayesian network (shown to the right). Each variable has two possible values, T (for true) and F (for false).
</p><p>The <a href="Joint_probability_distribution" title="Joint probability distribution">joint probability function</a> is, by the <a href="Chain_rule_of_probability" class="mw-redirect" title="Chain rule of probability">chain rule of probability</a>,
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(G,S,R)=\Pr(G\mid S,R)\Pr(S\mid R)\Pr(R)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>,</mo>
<mi>S</mi>
<mo>,</mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>∣<!-- ∣ --></mo>
<mi>S</mi>
<mo>,</mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>S</mi>
<mo>∣<!-- ∣ --></mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(G,S,R)=\Pr(G\mid S,R)\Pr(S\mid R)\Pr(R)}</annotation>
</semantics>
</math></span><img src="./8ebacd52e8162f5bc644c231eec89b4fa3aa6e41.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:43.271ex; height:2.843ex;" alt="{\displaystyle \Pr(G,S,R)=\Pr(G\mid S,R)\Pr(S\mid R)\Pr(R)}" loading="lazy"></span></dd></dl>
<p>where <i>G</i> = "Grass wet (true/false)", <i>S</i> = "Sprinkler turned on (true/false)", and <i>R</i> = "Raining (true/false)".
</p><p>The model can answer questions about the presence of a cause given the presence of an effect (so-called inverse probability) like "What is the probability that it is raining, given the grass is wet?" by using the <a href="Conditional_probability" title="Conditional probability">conditional probability</a> formula and summing over all <a href="Nuisance_variable" title="Nuisance variable">nuisance variables</a>:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(R=T\mid G=T)={\frac {\Pr(G=T,R=T)}{\Pr(G=T)}}={\frac {\sum _{x\in \{T,F\}}\Pr(G=T,S=x,R=T)}{\sum _{x,y\in \{T,F\}}\Pr(G=T,S=x,R=y)}}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo>=</mo>
<mi>T</mi>
<mo>∣<!-- ∣ --></mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mrow>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo>,</mo>
<mi>R</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
</mrow>
<mrow>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
</mrow>
</mfrac>
</mrow>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mrow>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>x</mi>
<mo>∈<!-- ∈ --></mo>
<mo fence="false" stretchy="false">{</mo>
<mi>T</mi>
<mo>,</mo>
<mi>F</mi>
<mo fence="false" stretchy="false">}</mo>
</mrow>
</munder>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo>,</mo>
<mi>S</mi>
<mo>=</mo>
<mi>x</mi>
<mo>,</mo>
<mi>R</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
</mrow>
<mrow>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>x</mi>
<mo>,</mo>
<mi>y</mi>
<mo>∈<!-- ∈ --></mo>
<mo fence="false" stretchy="false">{</mo>
<mi>T</mi>
<mo>,</mo>
<mi>F</mi>
<mo fence="false" stretchy="false">}</mo>
</mrow>
</munder>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo>,</mo>
<mi>S</mi>
<mo>=</mo>
<mi>x</mi>
<mo>,</mo>
<mi>R</mi>
<mo>=</mo>
<mi>y</mi>
<mo stretchy="false">)</mo>
</mrow>
</mfrac>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(R=T\mid G=T)={\frac {\Pr(G=T,R=T)}{\Pr(G=T)}}={\frac {\sum _{x\in \{T,F\}}\Pr(G=T,S=x,R=T)}{\sum _{x,y\in \{T,F\}}\Pr(G=T,S=x,R=y)}}}</annotation>
</semantics>
</math></span><img src="./eb7faabe239f8e620dab5f62d38f41193bad160c.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.171ex; width:81.32ex; height:7.509ex;" alt="{\displaystyle \Pr(R=T\mid G=T)={\frac {\Pr(G=T,R=T)}{\Pr(G=T)}}={\frac {\sum _{x\in \{T,F\}}\Pr(G=T,S=x,R=T)}{\sum _{x,y\in \{T,F\}}\Pr(G=T,S=x,R=y)}}}" loading="lazy"></span></dd></dl>
<p>Using the expansion for the joint probability function <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(G,S,R)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>,</mo>
<mi>S</mi>
<mo>,</mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(G,S,R)}</annotation>
</semantics>
</math></span><img src="./a60896f180f8cb353534d79a64ba0c9e8433c336.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:11.462ex; height:2.843ex;" alt="{\displaystyle \Pr(G,S,R)}" loading="lazy"></span> and the conditional probabilities from the <a href="Conditional_probability_table" title="Conditional probability table">conditional probability tables (CPTs)</a> stated in the diagram, one can evaluate each term in the sums in the numerator and denominator. For example,
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle {\begin{aligned}\Pr(G=T,S=T,R=T)&=\Pr(G=T\mid S=T,R=T)\Pr(S=T\mid R=T)\Pr(R=T)\\&=0.99\times 0.01\times 0.2\\&=0.00198.\end{aligned}}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mrow class="MJX-TeXAtom-ORD">
<mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true">
<mtr>
<mtd>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo>,</mo>
<mi>S</mi>
<mo>=</mo>
<mi>T</mi>
<mo>,</mo>
<mi>R</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
</mtd>
<mtd>
<mi></mi>
<mo>=</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo>∣<!-- ∣ --></mo>
<mi>S</mi>
<mo>=</mo>
<mi>T</mi>
<mo>,</mo>
<mi>R</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>S</mi>
<mo>=</mo>
<mi>T</mi>
<mo>∣<!-- ∣ --></mo>
<mi>R</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
</mtd>
</mtr>
<mtr>
<mtd></mtd>
<mtd>
<mi></mi>
<mo>=</mo>
<mn>0.99</mn>
<mo>×<!-- × --></mo>
<mn>0.01</mn>
<mo>×<!-- × --></mo>
<mn>0.2</mn>
</mtd>
</mtr>
<mtr>
<mtd></mtd>
<mtd>
<mi></mi>
<mo>=</mo>
<mn>0.00198.</mn>
</mtd>
</mtr>
</mtable>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle {\begin{aligned}\Pr(G=T,S=T,R=T)&=\Pr(G=T\mid S=T,R=T)\Pr(S=T\mid R=T)\Pr(R=T)\\&=0.99\times 0.01\times 0.2\\&=0.00198.\end{aligned}}}</annotation>
</semantics>
</math></span><img src="./a54d8f12a8e8ceaa849f0cf3b4a4c6b3426adafc.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.838ex; width:86.635ex; height:8.843ex;" alt="{\displaystyle {\begin{aligned}\Pr(G=T,S=T,R=T)&=\Pr(G=T\mid S=T,R=T)\Pr(S=T\mid R=T)\Pr(R=T)\\&=0.99\times 0.01\times 0.2\\&=0.00198.\end{aligned}}}" loading="lazy"></span></dd></dl>
<p>Then the numerical results (subscripted by the associated variable values) are
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(R=T\mid G=T)={\frac {0.00198_{TTT}+0.1584_{TFT}}{0.00198_{TTT}+0.288_{TTF}+0.1584_{TFT}+0.0_{TFF}}}={\frac {891}{2491}}\approx 35.77\%.}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo>=</mo>
<mi>T</mi>
<mo>∣<!-- ∣ --></mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mrow>
<msub>
<mn>0.00198</mn>
<mrow class="MJX-TeXAtom-ORD">
<mi>T</mi>
<mi>T</mi>
<mi>T</mi>
</mrow>
</msub>
<mo>+</mo>
<msub>
<mn>0.1584</mn>
<mrow class="MJX-TeXAtom-ORD">
<mi>T</mi>
<mi>F</mi>
<mi>T</mi>
</mrow>
</msub>
</mrow>
<mrow>
<msub>
<mn>0.00198</mn>
<mrow class="MJX-TeXAtom-ORD">
<mi>T</mi>
<mi>T</mi>
<mi>T</mi>
</mrow>
</msub>
<mo>+</mo>
<msub>
<mn>0.288</mn>
<mrow class="MJX-TeXAtom-ORD">
<mi>T</mi>
<mi>T</mi>
<mi>F</mi>
</mrow>
</msub>
<mo>+</mo>
<msub>
<mn>0.1584</mn>
<mrow class="MJX-TeXAtom-ORD">
<mi>T</mi>
<mi>F</mi>
<mi>T</mi>
</mrow>
</msub>
<mo>+</mo>
<msub>
<mn>0.0</mn>
<mrow class="MJX-TeXAtom-ORD">
<mi>T</mi>
<mi>F</mi>
<mi>F</mi>
</mrow>
</msub>
</mrow>
</mfrac>
</mrow>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mn>891</mn>
<mn>2491</mn>
</mfrac>
</mrow>
<mo>≈<!-- ≈ --></mo>
<mn>35.77</mn>
<mi mathvariant="normal">%<!-- % --></mi>
<mo>.</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(R=T\mid G=T)={\frac {0.00198_{TTT}+0.1584_{TFT}}{0.00198_{TTT}+0.288_{TTF}+0.1584_{TFT}+0.0_{TFF}}}={\frac {891}{2491}}\approx 35.77\%.}</annotation>
</semantics>
</math></span><img src="./e7b7d10d14d48f6bbbe766ce3e68216e01cf5829.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -2.171ex; width:88.777ex; height:5.509ex;" alt="{\displaystyle \Pr(R=T\mid G=T)={\frac {0.00198_{TTT}+0.1584_{TFT}}{0.00198_{TTT}+0.288_{TTF}+0.1584_{TFT}+0.0_{TFF}}}={\frac {891}{2491}}\approx 35.77\%.}" loading="lazy"></span></dd></dl>
<p>To answer an interventional question, such as "What is the probability that it would rain, given that we wet the grass?" the answer is governed by the post-intervention joint distribution function
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(S,R\mid {\text{do}}(G=T))=\Pr(S\mid R)\Pr(R)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>S</mi>
<mo>,</mo>
<mi>R</mi>
<mo>∣<!-- ∣ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mtext>do</mtext>
</mrow>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>S</mi>
<mo>∣<!-- ∣ --></mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(S,R\mid {\text{do}}(G=T))=\Pr(S\mid R)\Pr(R)}</annotation>
</semantics>
</math></span><img src="./b31a4e48eac8447fd1097a6be5ec510fcf7b4a0a.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:40.421ex; height:2.843ex;" alt="{\displaystyle \Pr(S,R\mid {\text{do}}(G=T))=\Pr(S\mid R)\Pr(R)}" loading="lazy"></span></dd></dl>
<p>obtained by removing the factor <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(G\mid S,R)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>∣<!-- ∣ --></mo>
<mi>S</mi>
<mo>,</mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(G\mid S,R)}</annotation>
</semantics>
</math></span><img src="./cbd899888f244eb37985d432ebe6db3fad449f80.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:12.365ex; height:2.843ex;" alt="{\displaystyle \Pr(G\mid S,R)}" loading="lazy"></span> from the pre-intervention distribution. The do operator forces the value of G to be true. The probability of rain is unaffected by the action:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(R\mid {\text{do}}(G=T))=\Pr(R).}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo>∣<!-- ∣ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mtext>do</mtext>
</mrow>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
<mo>.</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(R\mid {\text{do}}(G=T))=\Pr(R).}</annotation>
</semantics>
</math></span><img src="./0ba59bde63e1cb91dca3f046611c25a7dd75524c.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:28.644ex; height:2.843ex;" alt="{\displaystyle \Pr(R\mid {\text{do}}(G=T))=\Pr(R).}" loading="lazy"></span></dd></dl>
<p>To predict the impact of turning the sprinkler on:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(R,G\mid {\text{do}}(S=T))=\Pr(R)\Pr(G\mid R,S=T)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo>,</mo>
<mi>G</mi>
<mo>∣<!-- ∣ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mtext>do</mtext>
</mrow>
<mo stretchy="false">(</mo>
<mi>S</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>G</mi>
<mo>∣<!-- ∣ --></mo>
<mi>R</mi>
<mo>,</mo>
<mi>S</mi>
<mo>=</mo>
<mi>T</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(R,G\mid {\text{do}}(S=T))=\Pr(R)\Pr(G\mid R,S=T)}</annotation>
</semantics>
</math></span><img src="./b00365bf1cd38b5051a86ff9e72ced1509af038d.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:48.017ex; height:2.843ex;" alt="{\displaystyle \Pr(R,G\mid {\text{do}}(S=T))=\Pr(R)\Pr(G\mid R,S=T)}" loading="lazy"></span></dd></dl>
<p>with the term <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(S=T\mid R)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>S</mi>
<mo>=</mo>
<mi>T</mi>
<mo>∣<!-- ∣ --></mo>
<mi>R</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(S=T\mid R)}</annotation>
</semantics>
</math></span><img src="./cc6ac4d6d15fece56a8b824ddec919e4607d37cd.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:14.239ex; height:2.843ex;" alt="{\displaystyle \Pr(S=T\mid R)}" loading="lazy"></span> removed, showing that the action affects the grass but not the rain.
</p><p>These predictions may not be feasible given unobserved variables, as in most policy evaluation problems. The effect of the action <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle {\text{do}}(x)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mrow class="MJX-TeXAtom-ORD">
<mtext>do</mtext>
</mrow>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle {\text{do}}(x)}</annotation>
</semantics>
</math></span><img src="./47eff158e910cb801e05e0c0f353a8f4144f0925.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:5.594ex; height:2.843ex;" alt="{\displaystyle {\text{do}}(x)}" loading="lazy"></span> can still be predicted, however, whenever the back-door criterion is satisfied.<sup id="cite_ref-pearl2000_2-0" class="reference"><a href="#cite_note-pearl2000-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-3" class="reference"><a href="#cite_note-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> It states that, if a set <i>Z</i> of nodes can be observed that <a href="#d-separation"><i>d</i>-separates</a><sup id="cite_ref-4" class="reference"><a href="#cite_note-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> (or blocks) all back-door paths from <i>X</i> to <i>Y</i> then
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Pr(Y,Z\mid {\text{do}}(x))={\frac {\Pr(Y,Z,X=x)}{\Pr(X=x\mid Z)}}.}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>Y</mi>
<mo>,</mo>
<mi>Z</mi>
<mo>∣<!-- ∣ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mtext>do</mtext>
</mrow>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mrow>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>Y</mi>
<mo>,</mo>
<mi>Z</mi>
<mo>,</mo>
<mi>X</mi>
<mo>=</mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
</mrow>
<mrow>
<mo movablelimits="true" form="prefix">Pr</mo>
<mo stretchy="false">(</mo>
<mi>X</mi>
<mo>=</mo>
<mi>x</mi>
<mo>∣<!-- ∣ --></mo>
<mi>Z</mi>
<mo stretchy="false">)</mo>
</mrow>
</mfrac>
</mrow>
<mo>.</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Pr(Y,Z\mid {\text{do}}(x))={\frac {\Pr(Y,Z,X=x)}{\Pr(X=x\mid Z)}}.}</annotation>
</semantics>
</math></span><img src="./35cecd8237339fe753a6754961aae647e34f0a95.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -2.671ex; width:37.137ex; height:6.509ex;" alt="{\displaystyle \Pr(Y,Z\mid {\text{do}}(x))={\frac {\Pr(Y,Z,X=x)}{\Pr(X=x\mid Z)}}.}" loading="lazy"></span></dd></dl>
<p>A back-door path is one that ends with an arrow into <i>X</i>. Sets that satisfy the back-door criterion are called "sufficient" or "admissible." For example, the set <i>Z</i> = <i>R</i> is admissible for predicting the effect of <i>S</i> = <i>T</i> on <i>G</i>, because <i>R</i> <i>d</i>-separates the (only) back-door path <i>S</i> ← <i>R</i> → <i>G</i>. However, if <i>S</i> is not observed, no other set <i>d</i>-separates this path and the effect of turning the sprinkler on (<i>S</i> = <i>T</i>) on the grass (<i>G</i>) cannot be predicted from passive observations. In that case <i>P</i>(<i>G</i> | do(<i>S</i> = <i>T</i>)) is not "identified". This reflects the fact that, lacking interventional data, the observed dependence between <i>S</i> and <i>G</i> is due to a causal connection or is spurious
(apparent dependence arising from a common cause, <i>R</i>). (see <a href="Simpson's_paradox" title="Simpson's paradox">Simpson's paradox</a>)
</p><p>To determine whether a causal relation is identified from an arbitrary Bayesian network with unobserved variables, one can use the three rules of "<i>do</i>-calculus"<sup id="cite_ref-pearl2000_2-1" class="reference"><a href="#cite_note-pearl2000-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-pearl-r212_5-0" class="reference"><a href="#cite_note-pearl-r212-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> and test whether all <i>do</i> terms can be removed from the expression of that relation, thus confirming that the desired quantity is estimable from frequency data.<sup id="cite_ref-6" class="reference"><a href="#cite_note-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup>
</p><p>Using a Bayesian network can save considerable amounts of memory over exhaustive probability tables, if the dependencies in the joint distribution are sparse. For example, a naive way of storing the conditional probabilities of 10 two-valued variables as a table requires storage space for <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle 2^{10}=1024}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msup>
<mn>2</mn>
<mrow class="MJX-TeXAtom-ORD">
<mn>10</mn>
</mrow>
</msup>
<mo>=</mo>
<mn>1024</mn>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle 2^{10}=1024}</annotation>
</semantics>
</math></span><img src="./13588ba2bcf107e75098e0aac63663fa4f147e45.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:10.787ex; height:2.676ex;" alt="{\displaystyle 2^{10}=1024}" loading="lazy"></span> values. If no variable's local distribution depends on more than three parent variables, the Bayesian network representation stores at most <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle 10\cdot 2^{3}=80}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mn>10</mn>
<mo>⋅<!-- ⋅ --></mo>
<msup>
<mn>2</mn>
<mrow class="MJX-TeXAtom-ORD">
<mn>3</mn>
</mrow>
</msup>
<mo>=</mo>
<mn>80</mn>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle 10\cdot 2^{3}=80}</annotation>
</semantics>
</math></span><img src="./761979726d8db16a71c2caf0d8db784c8e8b319d.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:11.644ex; height:2.676ex;" alt="{\displaystyle 10\cdot 2^{3}=80}" loading="lazy"></span> values.
</p><p>One advantage of Bayesian networks is that it is intuitively easier for a human to understand (a sparse set of) direct dependencies and local distributions than complete joint distributions.
</p>
<div class="mw-heading mw-heading2"><h2 id="Inference_and_learning">Inference and learning</h2></div>
<p>Bayesian networks perform three main inference tasks:
</p>
<div class="mw-heading mw-heading3"><h3 id="Inferring_unobserved_variables">Inferring unobserved variables</h3></div>
<p>Because a Bayesian network is a complete model for its variables and their relationships, it can be used to answer probabilistic queries about them. For example, the network can be used to update knowledge of the state of a subset of variables when other variables (the <i>evidence</i> variables) are observed. This process of computing the <i>posterior</i> distribution of variables given evidence is called probabilistic inference. The posterior gives a universal <a href="Sufficient_statistic" title="Sufficient statistic">sufficient statistic</a> for detection applications, when choosing values for the variable subset that minimize some expected loss function, for instance the probability of decision error. A Bayesian network can thus be considered a mechanism for automatically applying <a href="Bayes'_theorem" title="Bayes' theorem">Bayes' theorem</a> to complex problems.
</p><p>The most common exact inference methods are: <a href="Variable_elimination" title="Variable elimination">variable elimination</a>, which eliminates (by integration or summation) the non-observed non-query variables one by one by distributing the sum over the product; <a href="Junction_tree_algorithm" title="Junction tree algorithm">clique tree propagation</a>, which caches the computation so that many variables can be queried at one time and new evidence can be propagated quickly; and recursive conditioning and AND/OR search, which allow for a <a href="Space%E2%80%93time_tradeoff" title="Space–time tradeoff">space–time tradeoff</a> and match the efficiency of variable elimination when enough space is used. All of these methods have complexity that is exponential in the network's <a href="Treewidth" title="Treewidth">treewidth</a>. The most common <a href="Approximate_inference" title="Approximate inference">approximate inference</a> algorithms are <a href="Importance_sampling" title="Importance sampling">importance sampling</a>, stochastic <a href="Markov_chain_Monte_Carlo" title="Markov chain Monte Carlo">MCMC</a> simulation, mini-bucket elimination, <a href="Loopy_belief_propagation" class="mw-redirect" title="Loopy belief propagation">loopy belief propagation</a>, <a href="Generalized_belief_propagation" class="mw-redirect" title="Generalized belief propagation">generalized belief propagation</a> and <a href="Variational_Bayes" class="mw-redirect" title="Variational Bayes">variational methods</a>.
</p>
<div class="mw-heading mw-heading3"><h3 id="Parameter_learning">Parameter learning</h3></div>
<p>In order to fully specify the Bayesian network and thus fully represent the <a href="Joint_probability_distribution" title="Joint probability distribution">joint probability distribution</a>, it is necessary to specify for each node <i>X</i> the probability distribution for <i>X</i> conditional upon <i>X</i><span class="nowrap" style="padding-left:0.1em;">'s</span> parents. The distribution of <i>X</i> conditional upon its parents may have any form. It is common to work with discrete or <a href="Normal_distribution" title="Normal distribution">Gaussian distributions</a> since that simplifies calculations. Sometimes only constraints on distribution are known; one can then use the <a href="Principle_of_maximum_entropy" title="Principle of maximum entropy">principle of maximum entropy</a> to determine a single distribution, the one with the greatest <a href="Information_entropy" class="mw-redirect" title="Information entropy">entropy</a> given the constraints. (Analogously, in the specific context of a <a href="Dynamic_Bayesian_network" title="Dynamic Bayesian network">dynamic Bayesian network</a>, the conditional distribution for the hidden state's temporal evolution is commonly specified to maximize the <a href="Entropy_rate" title="Entropy rate">entropy rate</a> of the implied stochastic process.)
</p><p>Often these conditional distributions include parameters that are unknown and must be estimated from data, e.g., via the <a href="Maximum_likelihood" class="mw-redirect" title="Maximum likelihood">maximum likelihood</a> approach. Direct maximization of the likelihood (or of the <a href="Posterior_probability" title="Posterior probability">posterior probability</a>) is often complex given unobserved variables. A classical approach to this problem is the <a href="Expectation-maximization_algorithm" class="mw-redirect" title="Expectation-maximization algorithm">expectation-maximization algorithm</a>, which alternates computing expected values of the unobserved variables conditional on observed data, with maximizing the complete likelihood (or posterior) assuming that previously computed expected values are correct. Under mild regularity conditions, this process converges on maximum likelihood (or maximum posterior) values for parameters.
</p><p>A more fully Bayesian approach to parameters is to treat them as additional unobserved variables and to compute a full posterior distribution over all nodes conditional upon observed data, then to integrate out the parameters. This approach can be expensive and lead to large dimension models, making classical parameter-setting approaches more tractable.
</p>
<div class="mw-heading mw-heading3"><h3 id="Structure_learning">Structure learning</h3></div>
<p>In the simplest case, a Bayesian network is specified by an expert and is then used to perform inference. In other applications, the task of defining the network is too complex for humans. In this case, the network structure and the parameters of the local distributions must be learned from data.
</p><p>Automatically learning the graph structure of a Bayesian network (BN) is a challenge pursued within <a href="Machine_learning" title="Machine learning">machine learning</a>. The basic idea goes back to a recovery algorithm developed by Rebane and <a href="Judea_Pearl" title="Judea Pearl">Pearl</a><sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup> and rests on the distinction between the three possible patterns allowed in a 3-node DAG:
</p>
<table class="wikitable">
<caption>Junction patterns
</caption>
<tbody><tr>
<th>Pattern
</th>
<th>Model
</th></tr>
<tr>
<td>Chain
</td>
<th><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X\rightarrow Y\rightarrow Z}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
<mo stretchy="false">→<!-- → --></mo>
<mi>Y</mi>
<mo stretchy="false">→<!-- → --></mo>
<mi>Z</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X\rightarrow Y\rightarrow Z}</annotation>
</semantics>
</math></span><img src="./ff5d092fb7193163f784b507c54b1c5781b2f453.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:12.662ex; height:2.176ex;" alt="{\displaystyle X\rightarrow Y\rightarrow Z}" loading="lazy"></span>
</th></tr>
<tr>
<td>Fork
</td>
<td><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X\leftarrow Y\rightarrow Z}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
<mo stretchy="false">←<!-- ← --></mo>
<mi>Y</mi>
<mo stretchy="false">→<!-- → --></mo>
<mi>Z</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X\leftarrow Y\rightarrow Z}</annotation>
</semantics>
</math></span><img src="./5d91a1d581c41a11fefca970fd25ab2419a920f4.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:12.662ex; height:2.176ex;" alt="{\displaystyle X\leftarrow Y\rightarrow Z}" loading="lazy"></span>
</td></tr>
<tr>
<td>Collider
</td>
<td><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X\rightarrow Y\leftarrow Z}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
<mo stretchy="false">→<!-- → --></mo>
<mi>Y</mi>
<mo stretchy="false">←<!-- ← --></mo>
<mi>Z</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X\rightarrow Y\leftarrow Z}</annotation>
</semantics>
</math></span><img src="./9b2c03efa1e7830a331531aca3e5b89a0fd9b763.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:12.662ex; height:2.176ex;" alt="{\displaystyle X\rightarrow Y\leftarrow Z}" loading="lazy"></span>
</td></tr></tbody></table>
<p>The first 2 represent the same dependencies (<span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X}</annotation>
</semantics>
</math></span><img src="./68baa052181f707c662844a465bfeeb135e82bab.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.98ex; height:2.176ex;" alt="{\displaystyle X}" loading="lazy"></span> and <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle Z}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>Z</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle Z}</annotation>
</semantics>
</math></span><img src="./1cc6b75e09a8aa3f04d8584b11db534f88fb56bd.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.68ex; height:2.176ex;" alt="{\displaystyle Z}" loading="lazy"></span> are independent given <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle Y}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>Y</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle Y}</annotation>
</semantics>
</math></span><img src="./961d67d6b454b4df2301ac571808a3538b3a6d3f.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.171ex; width:1.773ex; height:2.009ex;" alt="{\displaystyle Y}" loading="lazy"></span>) and are, therefore, indistinguishable. The collider, however, can be uniquely identified, since <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X}</annotation>
</semantics>
</math></span><img src="./68baa052181f707c662844a465bfeeb135e82bab.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.98ex; height:2.176ex;" alt="{\displaystyle X}" loading="lazy"></span> and <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle Z}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>Z</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle Z}</annotation>
</semantics>
</math></span><img src="./1cc6b75e09a8aa3f04d8584b11db534f88fb56bd.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.68ex; height:2.176ex;" alt="{\displaystyle Z}" loading="lazy"></span> are marginally independent and all other pairs are dependent. Thus, while the <i>skeletons</i> (the graphs stripped of arrows) of these three triplets are identical, the directionality of the arrows is partially identifiable. The same distinction applies when <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X}</annotation>
</semantics>
</math></span><img src="./68baa052181f707c662844a465bfeeb135e82bab.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.98ex; height:2.176ex;" alt="{\displaystyle X}" loading="lazy"></span> and <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle Z}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>Z</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle Z}</annotation>
</semantics>
</math></span><img src="./1cc6b75e09a8aa3f04d8584b11db534f88fb56bd.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.68ex; height:2.176ex;" alt="{\displaystyle Z}" loading="lazy"></span> have common parents, except that one must first condition on those parents. Algorithms have been developed to systematically determine the skeleton of the underlying graph and, then, orient all arrows whose directionality is dictated by the conditional independences observed.<sup id="cite_ref-pearl2000_2-2" class="reference"><a href="#cite_note-pearl2000-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-8" class="reference"><a href="#cite_note-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-10" class="reference"><a href="#cite_note-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup>
</p><p>An alternative method of structural learning uses optimization-based search. It requires a <a href="Scoring_function" class="mw-redirect" title="Scoring function">scoring function</a> and a search strategy. A common scoring function is <a href="Posterior_probability" title="Posterior probability">posterior probability</a> of the structure given the training data, like the <a href="Bayesian_information_criterion" title="Bayesian information criterion">BIC</a> or the BDeu. The time requirement of an <a href="Exhaustive_search" class="mw-redirect" title="Exhaustive search">exhaustive search</a> returning a structure that maximizes the score is <a href="Tetration" title="Tetration">superexponential</a> in the number of variables. A local search strategy makes incremental changes aimed at improving the score of the structure. A global search algorithm like <a href="Markov_chain_Monte_Carlo" title="Markov chain Monte Carlo">Markov chain Monte Carlo</a> can avoid getting trapped in <a href="Maxima_and_minima" class="mw-redirect" title="Maxima and minima">local minima</a>. Friedman et al.<sup id="cite_ref-11" class="reference"><a href="#cite_note-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-12" class="reference"><a href="#cite_note-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup> discuss using <a href="Mutual_information" title="Mutual information">mutual information</a> between variables and finding a structure that maximizes this. They do this by restricting the parent candidate set to <i>k</i> nodes and exhaustively searching therein.
</p><p>A particularly fast method for exact BN learning is to cast the problem as an optimization problem, and solve it using <a href="Integer_programming" title="Integer programming">integer programming</a>. Acyclicity constraints are added to the integer program (IP) during solving in the form of <a href="Cutting-plane_method" title="Cutting-plane method">cutting planes</a>.<sup id="cite_ref-13" class="reference"><a href="#cite_note-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup> Such method can handle problems with up to 100 variables.
</p><p>In order to deal with problems with thousands of variables, a different approach is necessary. One is to first sample one ordering, and then find the optimal BN structure with respect to that ordering. This implies working on the search space of the possible orderings, which is convenient as it is smaller than the space of network structures. Multiple orderings are then sampled and evaluated. This method has been proven to be the best available in literature when the number of variables is huge.<sup id="cite_ref-14" class="reference"><a href="#cite_note-14"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup>
</p><p>Another method consists of focusing on the sub-class of decomposable models, for which the <a href="Maximum_likelihood_estimate" class="mw-redirect" title="Maximum likelihood estimate">MLE</a> have a closed form. It is then possible to discover a consistent structure for hundreds of variables.<sup id="cite_ref-Petitjean_15-0" class="reference"><a href="#cite_note-Petitjean-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup>
</p><p>Learning Bayesian networks with bounded treewidth is necessary to allow exact, tractable inference, since the worst-case inference complexity is exponential in the treewidth k (under the exponential time hypothesis). Yet, as a global property of the graph, it considerably increases the difficulty of the learning process. In this context it is possible to use <a href="K-tree" title="K-tree">K-tree</a> for effective learning.<sup id="cite_ref-16" class="reference"><a href="#cite_note-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Statistical_introduction">Statistical introduction</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1236090951">
/* start https://en.wikipedia.org/ */
.mw-parser-output .hatnote{font-style:italic}.mw-parser-output div.hatnote{padding-left:1.6em;margin-bottom:0.5em}.mw-parser-output .hatnote i{font-style:normal}.mw-parser-output .hatnote+link+.hatnote{margin-top:-0.5em}@media print{body.ns-0 .mw-parser-output .hatnote{display:none!important}}
/* end https://en.wikipedia.org/ */
</style><div role="note" class="hatnote navigation-not-searchable">Main articles: <a href="Bayesian_statistics" title="Bayesian statistics">Bayesian statistics</a> and <a href="Multilevel_model" title="Multilevel model">Multilevel model</a></div>
<p>Given data <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle x\,\!}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>x</mi>
<mspace width="thinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle x\,\!}</annotation>
</semantics>
</math></span><img src="./ec21bb206c6c9f458130ab7ffddfe3fd8d0fa6bb.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; margin-right: -0.387ex; width:1.717ex; height:1.676ex;" alt="{\displaystyle x\,\!}" loading="lazy"></span> and parameter <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \theta }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>θ<!-- θ --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \theta }</annotation>
</semantics>
</math></span><img src="./6e5ab2664b422d53eb0c7df3b87e1360d75ad9af.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.09ex; height:2.176ex;" alt="{\displaystyle \theta }" loading="lazy"></span>, a simple <a href="Bayesian_statistics" title="Bayesian statistics">Bayesian analysis</a> starts with a <a href="Prior_probability" title="Prior probability">prior probability</a> (<i>prior</i>) <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(\theta )}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>θ<!-- θ --></mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(\theta )}</annotation>
</semantics>
</math></span><img src="./b66e40cc46404eb7159d0368dab9cfba9fd96471.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; margin-left: -0.089ex; width:4.159ex; height:2.843ex;" alt="{\displaystyle p(\theta )}" loading="lazy"></span> and <a href="Likelihood_function" title="Likelihood function">likelihood</a> <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(x\mid \theta )}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo>∣<!-- ∣ --></mo>
<mi>θ<!-- θ --></mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(x\mid \theta )}</annotation>
</semantics>
</math></span><img src="./0db2f0c3e32fa4373616fe08e36d9c405265331e.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; margin-left: -0.089ex; width:7.425ex; height:2.843ex;" alt="{\displaystyle p(x\mid \theta )}" loading="lazy"></span> to compute a <a href="Posterior_probability" title="Posterior probability">posterior probability</a> <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(\theta \mid x)\propto p(x\mid \theta )p(\theta )}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>θ<!-- θ --></mi>
<mo>∣<!-- ∣ --></mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo>∝<!-- ∝ --></mo>
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo>∣<!-- ∣ --></mo>
<mi>θ<!-- θ --></mi>
<mo stretchy="false">)</mo>
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>θ<!-- θ --></mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(\theta \mid x)\propto p(x\mid \theta )p(\theta )}</annotation>
</semantics>
</math></span><img src="./3a872cb198563edadb78ed2c6bb75a1749cb9047.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; margin-left: -0.089ex; width:21.929ex; height:2.843ex;" alt="{\displaystyle p(\theta \mid x)\propto p(x\mid \theta )p(\theta )}" loading="lazy"></span>.
</p><p>Often the prior on <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \theta }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>θ<!-- θ --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \theta }</annotation>
</semantics>
</math></span><img src="./6e5ab2664b422d53eb0c7df3b87e1360d75ad9af.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.09ex; height:2.176ex;" alt="{\displaystyle \theta }" loading="lazy"></span> depends in turn on other parameters <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \varphi }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>φ<!-- φ --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \varphi }</annotation>
</semantics>
</math></span><img src="./33ee699558d09cf9d653f6351f9fda0b2f4aaa3e.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:1.52ex; height:2.176ex;" alt="{\displaystyle \varphi }" loading="lazy"></span> that are not mentioned in the likelihood. So, the prior <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(\theta )}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>θ<!-- θ --></mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(\theta )}</annotation>
</semantics>
</math></span><img src="./b66e40cc46404eb7159d0368dab9cfba9fd96471.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; margin-left: -0.089ex; width:4.159ex; height:2.843ex;" alt="{\displaystyle p(\theta )}" loading="lazy"></span> must be replaced by a likelihood <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(\theta \mid \varphi )}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>θ<!-- θ --></mi>
<mo>∣<!-- ∣ --></mo>
<mi>φ<!-- φ --></mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(\theta \mid \varphi )}</annotation>
</semantics>
</math></span><img src="./e11cf423d84155bd96079646c3e0d63eb08db3d2.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; margin-left: -0.089ex; width:7.616ex; height:2.843ex;" alt="{\displaystyle p(\theta \mid \varphi )}" loading="lazy"></span>, and a prior <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(\varphi )}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>φ<!-- φ --></mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(\varphi )}</annotation>
</semantics>
</math></span><img src="./f75607255ff912e3800bcf2975185c44444a2863.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; margin-left: -0.089ex; width:4.588ex; height:2.843ex;" alt="{\displaystyle p(\varphi )}" loading="lazy"></span> on the newly introduced parameters <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \varphi }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>φ<!-- φ --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \varphi }</annotation>
</semantics>
</math></span><img src="./33ee699558d09cf9d653f6351f9fda0b2f4aaa3e.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:1.52ex; height:2.176ex;" alt="{\displaystyle \varphi }" loading="lazy"></span> is required, resulting in a posterior probability
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(\theta ,\varphi \mid x)\propto p(x\mid \theta )p(\theta \mid \varphi )p(\varphi ).}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>θ<!-- θ --></mi>
<mo>,</mo>
<mi>φ<!-- φ --></mi>
<mo>∣<!-- ∣ --></mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo>∝<!-- ∝ --></mo>
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo>∣<!-- ∣ --></mo>
<mi>θ<!-- θ --></mi>
<mo stretchy="false">)</mo>
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>θ<!-- θ --></mi>
<mo>∣<!-- ∣ --></mo>
<mi>φ<!-- φ --></mi>
<mo stretchy="false">)</mo>
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>φ<!-- φ --></mi>
<mo stretchy="false">)</mo>
<mo>.</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(\theta ,\varphi \mid x)\propto p(x\mid \theta )p(\theta \mid \varphi )p(\varphi ).}</annotation>
</semantics>
</math></span><img src="./461ec546c506493436a449dde12f9111190d7d9e.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; margin-left: -0.089ex; width:33.086ex; height:2.843ex;" alt="{\displaystyle p(\theta ,\varphi \mid x)\propto p(x\mid \theta )p(\theta \mid \varphi )p(\varphi ).}" loading="lazy"></span></dd></dl>
<p>This is the simplest example of a <a href="Bayesian_hierarchical_modeling" title="Bayesian hierarchical modeling"><i>hierarchical Bayes model</i></a>.
</p><p>The process may be repeated; for example, the parameters <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \varphi }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>φ<!-- φ --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \varphi }</annotation>
</semantics>
</math></span><img src="./33ee699558d09cf9d653f6351f9fda0b2f4aaa3e.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:1.52ex; height:2.176ex;" alt="{\displaystyle \varphi }" loading="lazy"></span> may depend in turn on additional parameters <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \psi \,\!}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>ψ<!-- ψ --></mi>
<mspace width="thinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \psi \,\!}</annotation>
</semantics>
</math></span><img src="./38353b935ec50c2f579b1cbad9ec1280247fefcf.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; margin-right: -0.387ex; width:1.9ex; height:2.509ex;" alt="{\displaystyle \psi \,\!}" loading="lazy"></span>, which require their own prior. Eventually the process must terminate, with priors that do not depend on unmentioned parameters.
</p>
<div class="mw-heading mw-heading3"><h3 id="Introductory_examples">Introductory examples</h3></div>
<p>Given the measured quantities <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle x_{1},\dots ,x_{n}\,\!}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>1</mn>
</mrow>
</msub>
<mo>,</mo>
<mo>…<!-- … --></mo>
<mo>,</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</msub>
<mspace width="thinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle x_{1},\dots ,x_{n}\,\!}</annotation>
</semantics>
</math></span><img src="./b2710506fdfe01459b37aecfd5c5846982128536.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; margin-right: -0.387ex; width:10.497ex; height:2.009ex;" alt="{\displaystyle x_{1},\dots ,x_{n}\,\!}" loading="lazy"></span>each with <a href="Normal_distribution" title="Normal distribution">normally distributed</a> errors of known <a href="Standard_deviation" title="Standard deviation">standard deviation</a> <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \sigma \,\!}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>σ<!-- σ --></mi>
<mspace width="thinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \sigma \,\!}</annotation>
</semantics>
</math></span><img src="./b645815a7c06785b9cc44600737d59624e4c75f7.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; margin-right: -0.387ex; width:1.717ex; height:1.676ex;" alt="{\displaystyle \sigma \,\!}" loading="lazy"></span>,
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle x_{i}\sim N(\theta _{i},\sigma ^{2})}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>∼<!-- ∼ --></mo>
<mi>N</mi>
<mo stretchy="false">(</mo>
<msub>
<mi>θ<!-- θ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>,</mo>
<msup>
<mi>σ<!-- σ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>2</mn>
</mrow>
</msup>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle x_{i}\sim N(\theta _{i},\sigma ^{2})}</annotation>
</semantics>
</math></span><img src="./38a1dabd34519ee9e7d386fb9630a740fc57e44c.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:14.41ex; height:3.176ex;" alt="{\displaystyle x_{i}\sim N(\theta _{i},\sigma ^{2})}" loading="lazy"></span></dd></dl>
<p>Suppose we are interested in estimating the <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \theta _{i}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>θ<!-- θ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \theta _{i}}</annotation>
</semantics>
</math></span><img src="./302b19204ed378e99ff4575341a67eebdbe5a555.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:1.89ex; height:2.509ex;" alt="{\displaystyle \theta _{i}}" loading="lazy"></span>. An approach would be to estimate the <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \theta _{i}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>θ<!-- θ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \theta _{i}}</annotation>
</semantics>
</math></span><img src="./302b19204ed378e99ff4575341a67eebdbe5a555.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:1.89ex; height:2.509ex;" alt="{\displaystyle \theta _{i}}" loading="lazy"></span> using a <a href="Maximum_likelihood" class="mw-redirect" title="Maximum likelihood">maximum likelihood</a> approach; since the observations are independent, the likelihood factorizes and the maximum likelihood estimate is simply
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \theta _{i}=x_{i}.}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>θ<!-- θ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>.</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \theta _{i}=x_{i}.}</annotation>
</semantics>
</math></span><img src="./8d5752b677e8949e30c16e26b1c6a32cd24e2ac9.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:7.765ex; height:2.509ex;" alt="{\displaystyle \theta _{i}=x_{i}.}" loading="lazy"></span></dd></dl>
<p>However, if the quantities are related, so that for example the individual <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \theta _{i}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>θ<!-- θ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \theta _{i}}</annotation>
</semantics>
</math></span><img src="./302b19204ed378e99ff4575341a67eebdbe5a555.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:1.89ex; height:2.509ex;" alt="{\displaystyle \theta _{i}}" loading="lazy"></span>have themselves been drawn from an underlying distribution, then this relationship destroys the independence and suggests a more complex model, e.g.,
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle x_{i}\sim N(\theta _{i},\sigma ^{2}),}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>∼<!-- ∼ --></mo>
<mi>N</mi>
<mo stretchy="false">(</mo>
<msub>
<mi>θ<!-- θ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>,</mo>
<msup>
<mi>σ<!-- σ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>2</mn>
</mrow>
</msup>
<mo stretchy="false">)</mo>
<mo>,</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle x_{i}\sim N(\theta _{i},\sigma ^{2}),}</annotation>
</semantics>
</math></span><img src="./d7d4c771745ddb4f838c9d2f0089e7f7ba391f40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:15.056ex; height:3.176ex;" alt="{\displaystyle x_{i}\sim N(\theta _{i},\sigma ^{2}),}" loading="lazy"></span></dd>
<dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \theta _{i}\sim N(\varphi ,\tau ^{2}),}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>θ<!-- θ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>∼<!-- ∼ --></mo>
<mi>N</mi>
<mo stretchy="false">(</mo>
<mi>φ<!-- φ --></mi>
<mo>,</mo>
<msup>
<mi>τ<!-- τ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>2</mn>
</mrow>
</msup>
<mo stretchy="false">)</mo>
<mo>,</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \theta _{i}\sim N(\varphi ,\tau ^{2}),}</annotation>
</semantics>
</math></span><img src="./d4e32a8acdde8f07ee86016aeda6cb32ffdb5cd1.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:14.374ex; height:3.176ex;" alt="{\displaystyle \theta _{i}\sim N(\varphi ,\tau ^{2}),}" loading="lazy"></span></dd></dl>
<p>with <a href="Improper_prior" class="mw-redirect" title="Improper prior">improper priors</a> <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \varphi \sim {\text{flat}}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>φ<!-- φ --></mi>
<mo>∼<!-- ∼ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mtext>flat</mtext>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \varphi \sim {\text{flat}}}</annotation>
</semantics>
</math></span><img src="./5ca7bbe9400efc1f3539a1fe9055a4af07c86ddc.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:8.044ex; height:2.676ex;" alt="{\displaystyle \varphi \sim {\text{flat}}}" loading="lazy"></span>, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \tau \sim {\text{flat}}\in (0,\infty )}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>τ<!-- τ --></mi>
<mo>∼<!-- ∼ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mtext>flat</mtext>
</mrow>
<mo>∈<!-- ∈ --></mo>
<mo stretchy="false">(</mo>
<mn>0</mn>
<mo>,</mo>
<mi mathvariant="normal">∞<!-- ∞ --></mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \tau \sim {\text{flat}}\in (0,\infty )}</annotation>
</semantics>
</math></span><img src="./58049aa227de80c620c7e1b1546ae6204ae862ab.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:16.896ex; height:2.843ex;" alt="{\displaystyle \tau \sim {\text{flat}}\in (0,\infty )}" loading="lazy"></span>. When <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle n\geq 3}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>n</mi>
<mo>≥<!-- ≥ --></mo>
<mn>3</mn>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle n\geq 3}</annotation>
</semantics>
</math></span><img src="./73136e4a27fe39c123d16a7808e76d3162ce42bb.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.505ex; width:5.656ex; height:2.343ex;" alt="{\displaystyle n\geq 3}" loading="lazy"></span>, this is an <i>identified model</i> (i.e. there exists a unique solution for the model's parameters), and the posterior distributions of the individual <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \theta _{i}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>θ<!-- θ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \theta _{i}}</annotation>
</semantics>
</math></span><img src="./302b19204ed378e99ff4575341a67eebdbe5a555.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:1.89ex; height:2.509ex;" alt="{\displaystyle \theta _{i}}" loading="lazy"></span> will tend to move, or <i><a href="Shrinkage_estimator" class="mw-redirect" title="Shrinkage estimator">shrink</a></i> away from the maximum likelihood estimates towards their common mean. This <i>shrinkage</i> is a typical behavior in hierarchical Bayes models.
</p>
<div class="mw-heading mw-heading3"><h3 id="Restrictions_on_priors">Restrictions on priors</h3></div>
<p>Some care is needed when choosing priors in a hierarchical model, particularly on scale variables at higher levels of the hierarchy such as the variable <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \tau \,\!}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>τ<!-- τ --></mi>
<mspace width="thinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \tau \,\!}</annotation>
</semantics>
</math></span><img src="./1a1e3c1c15f9f70a61aafed2b37b24ec3b810e36.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; margin-right: -0.387ex; width:1.589ex; height:1.676ex;" alt="{\displaystyle \tau \,\!}" loading="lazy"></span> in the example. The usual priors such as the <a href="Jeffreys_prior" title="Jeffreys prior">Jeffreys prior</a> often do not work, because the posterior distribution will not be normalizable and estimates made by minimizing the <a href="Loss_function#Expected_loss" title="Loss function">expected loss</a> will be <a href="Admissible_decision_rule" title="Admissible decision rule">inadmissible</a>.
</p>
<div class="mw-heading mw-heading2"><h2 id="Definitions_and_concepts">Definitions and concepts</h2></div>
<div role="note" class="hatnote navigation-not-searchable">See also: <a href="Glossary_of_graph_theory#Directed_acyclic_graphs" title="Glossary of graph theory">Glossary of graph theory § Directed acyclic graphs</a></div>
<p>Several equivalent definitions of a Bayesian network have been offered. For the following, let <i>G</i> = (<i>V</i>,<i>E</i>) be a <a href="Directed_acyclic_graph" title="Directed acyclic graph">directed acyclic graph</a> (DAG) and let <i>X</i> = (<i>X</i><sub><i>v</i></sub>), <i>v</i> ∈ <i>V</i> be a set of <a href="Random_variable" title="Random variable">random variables</a> indexed by <i>V</i>.
</p>
<div class="mw-heading mw-heading3"><h3 id="Factorization_definition">Factorization definition</h3></div>
<p><i>X</i> is a Bayesian network with respect to <i>G</i> if its joint <a href="Probability_density_function" title="Probability density function">probability density function</a> (with respect to a <a href="Product_measure" title="Product measure">product measure</a>) can be written as a product of the individual density functions, conditional on their parent variables:<sup id="cite_ref-FOOTNOTERussellNorvig2003496_17-0" class="reference"><a href="#cite_note-FOOTNOTERussellNorvig2003496-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup>
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(x)=\prod _{v\in V}p\left(x_{v}\,{\big |}\,x_{\operatorname {pa} (v)}\right)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<munder>
<mo>∏<!-- ∏ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
<mo>∈<!-- ∈ --></mo>
<mi>V</mi>
</mrow>
</munder>
<mi>p</mi>
<mrow>
<mo>(</mo>
<mrow>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mspace width="thinmathspace"></mspace>
<mrow class="MJX-TeXAtom-ORD">
<mrow class="MJX-TeXAtom-ORD">
<mo maxsize="1.2em" minsize="1.2em">|</mo>
</mrow>
</mrow>
<mspace width="thinmathspace"></mspace>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>pa</mi>
<mo><!-- --></mo>
<mo stretchy="false">(</mo>
<mi>v</mi>
<mo stretchy="false">)</mo>
</mrow>
</msub>
</mrow>
<mo>)</mo>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(x)=\prod _{v\in V}p\left(x_{v}\,{\big |}\,x_{\operatorname {pa} (v)}\right)}</annotation>
</semantics>
</math></span><img src="./d98087c49ceff495def4e7f24e686ad9a953117c.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.171ex; margin-left: -0.089ex; width:23.882ex; height:5.676ex;" alt="{\displaystyle p(x)=\prod _{v\in V}p\left(x_{v}\,{\big |}\,x_{\operatorname {pa} (v)}\right)}" loading="lazy"></span></dd></dl>
<p>where pa(<i>v</i>) is the set of parents of <i>v</i> (i.e. those vertices pointing directly to <i>v</i> via a single edge).
</p><p>For any set of random variables, the probability of any member of a <a href="Joint_distribution" class="mw-redirect" title="Joint distribution">joint distribution</a> can be calculated from conditional probabilities using the <a href="Chain_rule_(probability)" title="Chain rule (probability)">chain rule</a> (given a <a href="Topological_ordering" class="mw-redirect" title="Topological ordering">topological ordering</a> of <i>X</i>) as follows:<sup id="cite_ref-FOOTNOTERussellNorvig2003496_17-1" class="reference"><a href="#cite_note-FOOTNOTERussellNorvig2003496-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup>
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \operatorname {P} (X_{1}=x_{1},\ldots ,X_{n}=x_{n})=\prod _{v=1}^{n}\operatorname {P} \left(X_{v}=x_{v}\mid X_{v+1}=x_{v+1},\ldots ,X_{n}=x_{n}\right)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi mathvariant="normal">P</mi>
<mo><!-- --></mo>
<mo stretchy="false">(</mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>1</mn>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>1</mn>
</mrow>
</msub>
<mo>,</mo>
<mo>…<!-- … --></mo>
<mo>,</mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</msub>
<mo stretchy="false">)</mo>
<mo>=</mo>
<munderover>
<mo>∏<!-- ∏ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
<mo>=</mo>
<mn>1</mn>
</mrow>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</munderover>
<mi mathvariant="normal">P</mi>
<mo><!-- --></mo>
<mrow>
<mo>(</mo>
<mrow>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>∣<!-- ∣ --></mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
<mo>+</mo>
<mn>1</mn>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
<mo>+</mo>
<mn>1</mn>
</mrow>
</msub>
<mo>,</mo>
<mo>…<!-- … --></mo>
<mo>,</mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</msub>
</mrow>
<mo>)</mo>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \operatorname {P} (X_{1}=x_{1},\ldots ,X_{n}=x_{n})=\prod _{v=1}^{n}\operatorname {P} \left(X_{v}=x_{v}\mid X_{v+1}=x_{v+1},\ldots ,X_{n}=x_{n}\right)}</annotation>
</semantics>
</math></span><img src="./adda3a3b3bd6a521a9923930bda814d5a129f157.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.005ex; width:72.597ex; height:6.843ex;" alt="{\displaystyle \operatorname {P} (X_{1}=x_{1},\ldots ,X_{n}=x_{n})=\prod _{v=1}^{n}\operatorname {P} \left(X_{v}=x_{v}\mid X_{v+1}=x_{v+1},\ldots ,X_{n}=x_{n}\right)}" loading="lazy"></span></dd></dl>
<p>Using the definition above, this can be written as:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \operatorname {P} (X_{1}=x_{1},\ldots ,X_{n}=x_{n})=\prod _{v=1}^{n}\operatorname {P} (X_{v}=x_{v}\mid X_{j}=x_{j}{\text{ for each }}X_{j}\,{\text{ that is a parent of }}X_{v}\,)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi mathvariant="normal">P</mi>
<mo><!-- --></mo>
<mo stretchy="false">(</mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>1</mn>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>1</mn>
</mrow>
</msub>
<mo>,</mo>
<mo>…<!-- … --></mo>
<mo>,</mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</msub>
<mo stretchy="false">)</mo>
<mo>=</mo>
<munderover>
<mo>∏<!-- ∏ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
<mo>=</mo>
<mn>1</mn>
</mrow>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</munderover>
<mi mathvariant="normal">P</mi>
<mo><!-- --></mo>
<mo stretchy="false">(</mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>∣<!-- ∣ --></mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mtext> for each </mtext>
</mrow>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mspace width="thinmathspace"></mspace>
<mrow class="MJX-TeXAtom-ORD">
<mtext> that is a parent of </mtext>
</mrow>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mspace width="thinmathspace"></mspace>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \operatorname {P} (X_{1}=x_{1},\ldots ,X_{n}=x_{n})=\prod _{v=1}^{n}\operatorname {P} (X_{v}=x_{v}\mid X_{j}=x_{j}{\text{ for each }}X_{j}\,{\text{ that is a parent of }}X_{v}\,)}</annotation>
</semantics>
</math></span><img src="./705ed59bfbeda83a3190c26437977d43a735a302.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.005ex; width:88.742ex; height:6.843ex;" alt="{\displaystyle \operatorname {P} (X_{1}=x_{1},\ldots ,X_{n}=x_{n})=\prod _{v=1}^{n}\operatorname {P} (X_{v}=x_{v}\mid X_{j}=x_{j}{\text{ for each }}X_{j}\,{\text{ that is a parent of }}X_{v}\,)}" loading="lazy"></span></dd></dl>
<p>The difference between the two expressions is the <a href="Conditional_independence" title="Conditional independence">conditional independence</a> of the variables from any of their non-descendants, given the values of their parent variables.
</p>
<div class="mw-heading mw-heading3"><h3 id="Local_Markov_property">Local Markov property</h3></div>
<p><i>X</i> is a Bayesian network with respect to <i>G</i> if it satisfies the <i>local Markov property</i>: each variable is <a href="Conditional_independence" title="Conditional independence">conditionally independent</a> of its non-descendants given its parent variables:<sup id="cite_ref-FOOTNOTERussellNorvig2003499_18-0" class="reference"><a href="#cite_note-FOOTNOTERussellNorvig2003499-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup>
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X_{v}\perp \!\!\!\perp X_{V\,\smallsetminus \,\operatorname {de} (v)}\mid X_{\operatorname {pa} (v)}\quad {\text{for all }}v\in V}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>⊥<!-- ⊥ --></mo>
<mspace width="negativethinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
<mo>⊥<!-- ⊥ --></mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>V</mi>
<mspace width="thinmathspace"></mspace>
<mo>∖<!-- ∖ --></mo>
<mspace width="thinmathspace"></mspace>
<mi>de</mi>
<mo><!-- --></mo>
<mo stretchy="false">(</mo>
<mi>v</mi>
<mo stretchy="false">)</mo>
</mrow>
</msub>
<mo>∣<!-- ∣ --></mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>pa</mi>
<mo><!-- --></mo>
<mo stretchy="false">(</mo>
<mi>v</mi>
<mo stretchy="false">)</mo>
</mrow>
</msub>
<mspace width="1em"></mspace>
<mrow class="MJX-TeXAtom-ORD">
<mtext>for all </mtext>
</mrow>
<mi>v</mi>
<mo>∈<!-- ∈ --></mo>
<mi>V</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X_{v}\perp \!\!\!\perp X_{V\,\smallsetminus \,\operatorname {de} (v)}\mid X_{\operatorname {pa} (v)}\quad {\text{for all }}v\in V}</annotation>
</semantics>
</math></span><img src="./adaf6e6e176a97f1d032671c7c1fe791c66b7a65.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -1.171ex; width:38.281ex; height:3.176ex;" alt="{\displaystyle X_{v}\perp \!\!\!\perp X_{V\,\smallsetminus \,\operatorname {de} (v)}\mid X_{\operatorname {pa} (v)}\quad {\text{for all }}v\in V}" loading="lazy"></span></dd></dl>
<p>where de(<i>v</i>) is the set of descendants and <i>V</i> \ de(<i>v</i>) is the set of non-descendants of <i>v</i>.
</p><p>This can be expressed in terms similar to the first definition, as
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle {\begin{aligned}&\operatorname {P} (X_{v}=x_{v}\mid X_{i}=x_{i}{\text{ for each }}X_{i}{\text{ that is not a descendant of }}X_{v}\,)\\[6pt]={}&P(X_{v}=x_{v}\mid X_{j}=x_{j}{\text{ for each }}X_{j}{\text{ that is a parent of }}X_{v}\,)\end{aligned}}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mrow class="MJX-TeXAtom-ORD">
<mtable columnalign="right left right left right left right left right left right left" rowspacing="0.9em 0.3em" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true">
<mtr>
<mtd></mtd>
<mtd>
<mi mathvariant="normal">P</mi>
<mo><!-- --></mo>
<mo stretchy="false">(</mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>∣<!-- ∣ --></mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mtext> for each </mtext>
</mrow>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mtext> that is not a descendant of </mtext>
</mrow>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mspace width="thinmathspace"></mspace>
<mo stretchy="false">)</mo>
</mtd>
</mtr>
<mtr>
<mtd>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
</mrow>
</mtd>
<mtd>
<mi>P</mi>
<mo stretchy="false">(</mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>∣<!-- ∣ --></mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mtext> for each </mtext>
</mrow>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mtext> that is a parent of </mtext>
</mrow>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mspace width="thinmathspace"></mspace>
<mo stretchy="false">)</mo>
</mtd>
</mtr>
</mtable>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle {\begin{aligned}&\operatorname {P} (X_{v}=x_{v}\mid X_{i}=x_{i}{\text{ for each }}X_{i}{\text{ that is not a descendant of }}X_{v}\,)\\[6pt]={}&P(X_{v}=x_{v}\mid X_{j}=x_{j}{\text{ for each }}X_{j}{\text{ that is a parent of }}X_{v}\,)\end{aligned}}}</annotation>
</semantics>
</math></span><img src="./f8f18a3121835ff1e58881ebfc0d5f8ef95b783d.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.171ex; width:67.549ex; height:7.509ex;" alt="{\displaystyle {\begin{aligned}&\operatorname {P} (X_{v}=x_{v}\mid X_{i}=x_{i}{\text{ for each }}X_{i}{\text{ that is not a descendant of }}X_{v}\,)\\[6pt]={}&P(X_{v}=x_{v}\mid X_{j}=x_{j}{\text{ for each }}X_{j}{\text{ that is a parent of }}X_{v}\,)\end{aligned}}}" loading="lazy"></span></dd></dl>
<p>The set of parents is a subset of the set of non-descendants because the graph is <a href="Cycle_(graph_theory)" title="Cycle (graph theory)">acyclic</a>.
</p>
<div class="mw-heading mw-heading3"><h3 id="Marginal_independence_structure">Marginal independence structure</h3></div>
<p>In general, learning a Bayesian network from data is known to be <a href="NP-hardness" title="NP-hardness">NP-hard</a>.<sup id="cite_ref-19" class="reference"><a href="#cite_note-19"><span class="cite-bracket">[</span>19<span class="cite-bracket">]</span></a></sup> This is due in part to the <a href="Combinatorial_explosion" title="Combinatorial explosion">combinatorial explosion</a> of <a href="Directed_acyclic_graph#Combinatorial_enumeration" title="Directed acyclic graph">enumerating DAGs</a> as the number of variables increases. Nevertheless, insights about an underlying Bayesian network can be learned from data in polynomial time by focusing on its marginal independence structure:<sup id="cite_ref-20" class="reference"><a href="#cite_note-20"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup> while the conditional independence statements of a distribution modeled by a Bayesian network are encoded by a DAG (according to the factorization and Markov properties above), its marginal independence statements—the conditional independence statements in which the conditioning set is empty—are encoded by a <a href="Graph_(discrete_mathematics)" title="Graph (discrete mathematics)">simple undirected graph</a> with special properties such as equal <a href="Intersection_number_(graph_theory)" title="Intersection number (graph theory)">intersection</a> and <a href="Independent_set_(graph_theory)" title="Independent set (graph theory)">independence numbers</a>.
</p>
<div class="mw-heading mw-heading3"><h3 id="Developing_Bayesian_networks">Developing Bayesian networks</h3></div>
<p>Developing a Bayesian network often begins with creating a DAG <i>G</i> such that <i>X</i> satisfies the local Markov property with respect to <i>G</i>. Sometimes this is a <a href="Causal_graph" title="Causal graph">causal</a> DAG. The conditional probability distributions of each variable given its parents in <i>G</i> are assessed. In many cases, in particular in the case where the variables are discrete, if the joint distribution of <i>X</i> is the product of these conditional distributions, then <i>X</i> is a Bayesian network with respect to <i>G</i>.<sup id="cite_ref-21" class="reference"><a href="#cite_note-21"><span class="cite-bracket">[</span>21<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Markov_blanket">Markov blanket</h3></div>
<p>The <a href="Markov_blanket" title="Markov blanket">Markov blanket</a> of a node is the set of nodes consisting of its parents, its children, and any other parents of its children. The Markov blanket renders the node independent of the rest of the network; the joint distribution of the variables in the Markov blanket of a node is sufficient knowledge for calculating the distribution of the node. <i>X</i> is a Bayesian network with respect to <i>G</i> if every node is conditionally independent of all other nodes in the network, given its <a href="Markov_blanket" title="Markov blanket">Markov blanket</a>.<sup id="cite_ref-FOOTNOTERussellNorvig2003499_18-1" class="reference"><a href="#cite_note-FOOTNOTERussellNorvig2003499-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading4"><h4 id="d-separation"><i>d</i>-separation</h4></div>
<p>This definition can be made more general by defining the "d"-separation of two nodes, where d stands for directional.<sup id="cite_ref-pearl2000_2-3" class="reference"><a href="#cite_note-pearl2000-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> We first define the "d"-separation of a trail and then we will define the "d"-separation of two nodes in terms of that.
</p><p>Let <i>P</i> be a trail from node <i>u</i> to <i>v</i>. A trail is a loop-free, undirected (i.e. all edge directions are ignored) path between two nodes. Then <i>P</i> is said to be <i>d</i>-separated by a set of nodes <i>Z</i> if any of the following conditions holds:
</p>
<ul><li><i>P</i> contains (but does not need to be entirely) a directed chain, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle u\cdots \leftarrow m\leftarrow \cdots v}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>u</mi>
<mo>⋯<!-- ⋯ --></mo>
<mo stretchy="false">←<!-- ← --></mo>
<mi>m</mi>
<mo stretchy="false">←<!-- ← --></mo>
<mo>⋯<!-- ⋯ --></mo>
<mi>v</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle u\cdots \leftarrow m\leftarrow \cdots v}</annotation>
</semantics>
</math></span><img src="./f9076c698cff90cd6a30a8f5a329b01a96bef635.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:17.947ex; height:1.843ex;" alt="{\displaystyle u\cdots \leftarrow m\leftarrow \cdots v}" loading="lazy"></span> or <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle u\cdots \rightarrow m\rightarrow \cdots v}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>u</mi>
<mo>⋯<!-- ⋯ --></mo>
<mo stretchy="false">→<!-- → --></mo>
<mi>m</mi>
<mo stretchy="false">→<!-- → --></mo>
<mo>⋯<!-- ⋯ --></mo>
<mi>v</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle u\cdots \rightarrow m\rightarrow \cdots v}</annotation>
</semantics>
</math></span><img src="./ff581c919fd3927f7c4b1beae79084a21e027a5a.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:17.947ex; height:1.843ex;" alt="{\displaystyle u\cdots \rightarrow m\rightarrow \cdots v}" loading="lazy"></span>, such that the middle node <i>m</i> is in <i>Z</i>,</li>
<li><i>P</i> contains a fork, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle u\cdots \leftarrow m\rightarrow \cdots v}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>u</mi>
<mo>⋯<!-- ⋯ --></mo>
<mo stretchy="false">←<!-- ← --></mo>
<mi>m</mi>
<mo stretchy="false">→<!-- → --></mo>
<mo>⋯<!-- ⋯ --></mo>
<mi>v</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle u\cdots \leftarrow m\rightarrow \cdots v}</annotation>
</semantics>
</math></span><img src="./aa01ad04fd6f6cecfda4bbefa0842b1822206f6d.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:17.947ex; height:1.843ex;" alt="{\displaystyle u\cdots \leftarrow m\rightarrow \cdots v}" loading="lazy"></span>, such that the middle node <i>m</i> is in <i>Z</i>, or</li>
<li><i>P</i> contains an inverted fork (or collider), <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle u\cdots \rightarrow m\leftarrow \cdots v}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>u</mi>
<mo>⋯<!-- ⋯ --></mo>
<mo stretchy="false">→<!-- → --></mo>
<mi>m</mi>
<mo stretchy="false">←<!-- ← --></mo>
<mo>⋯<!-- ⋯ --></mo>
<mi>v</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle u\cdots \rightarrow m\leftarrow \cdots v}</annotation>
</semantics>
</math></span><img src="./80b372b0185d692cd5e0a72f1c365551ddfb9b66.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:17.947ex; height:1.843ex;" alt="{\displaystyle u\cdots \rightarrow m\leftarrow \cdots v}" loading="lazy"></span>, such that the middle node <i>m</i> is not in <i>Z</i> and no descendant of <i>m</i> is in <i>Z</i>.</li></ul>
<p>The nodes <i>u</i> and <i>v</i> are <i>d</i>-separated by <i>Z</i> if all trails between them are <i>d</i>-separated. If <i>u</i> and <i>v</i> are not d-separated, they are d-connected.
</p><p><i>X</i> is a Bayesian network with respect to <i>G</i> if, for any two nodes <i>u</i>, <i>v</i>:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X_{u}\perp \!\!\!\perp X_{v}\mid X_{Z}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>u</mi>
</mrow>
</msub>
<mo>⊥<!-- ⊥ --></mo>
<mspace width="negativethinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
<mspace width="negativethinmathspace"></mspace>
<mo>⊥<!-- ⊥ --></mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>v</mi>
</mrow>
</msub>
<mo>∣<!-- ∣ --></mo>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>Z</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X_{u}\perp \!\!\!\perp X_{v}\mid X_{Z}}</annotation>
</semantics>
</math></span><img src="./15fa99e29ac7e72665f80dda38845df22bd3aed6.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:15.078ex; height:2.843ex;" alt="{\displaystyle X_{u}\perp \!\!\!\perp X_{v}\mid X_{Z}}" loading="lazy"></span></dd></dl>
<p>where <i>Z</i> is a set which <i>d</i>-separates <i>u</i> and <i>v</i>. (The <a href="Markov_blanket" title="Markov blanket">Markov blanket</a> is the minimal set of nodes which <i>d</i>-separates node <i>v</i> from all other nodes.)
</p>
<div class="mw-heading mw-heading3"><h3 id="Causal_networks">Causal networks</h3></div>
<p>Although Bayesian networks are often used to represent <a href="Causality" title="Causality">causal</a> relationships, this need not be the case: a directed edge from <i>u</i> to <i>v</i> does not require that <i>X<sub>v</sub></i> be causally dependent on <i>X<sub>u</sub></i>. This is demonstrated by the fact that Bayesian networks on the graphs:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle a\rightarrow b\rightarrow c\qquad {\text{and}}\qquad a\leftarrow b\leftarrow c}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>a</mi>
<mo stretchy="false">→<!-- → --></mo>
<mi>b</mi>
<mo stretchy="false">→<!-- → --></mo>
<mi>c</mi>
<mspace width="2em"></mspace>
<mrow class="MJX-TeXAtom-ORD">
<mtext>and</mtext>
</mrow>
<mspace width="2em"></mspace>
<mi>a</mi>
<mo stretchy="false">←<!-- ← --></mo>
<mi>b</mi>
<mo stretchy="false">←<!-- ← --></mo>
<mi>c</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle a\rightarrow b\rightarrow c\qquad {\text{and}}\qquad a\leftarrow b\leftarrow c}</annotation>
</semantics>
</math></span><img src="./237590f76460268148ca8cc06da9f2b1181160c6.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:33.963ex; height:2.176ex;" alt="{\displaystyle a\rightarrow b\rightarrow c\qquad {\text{and}}\qquad a\leftarrow b\leftarrow c}" loading="lazy"></span></dd></dl>
<p>are equivalent: that is they impose exactly the same conditional independence requirements.
</p><p>A causal network is a Bayesian network with the requirement that the relationships be causal. The additional semantics of causal networks specify that if a node <i>X</i> is actively caused to be in a given state <i>x</i> (an action written as do(<i>X</i> = <i>x</i>)), then the probability density function changes to that of the network obtained by cutting the links from the parents of <i>X</i> to <i>X</i>, and setting <i>X</i> to the caused value <i>x</i>.<sup id="cite_ref-pearl2000_2-4" class="reference"><a href="#cite_note-pearl2000-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> Using these semantics, the impact of external interventions from data obtained prior to intervention can be predicted.
</p>
<div class="mw-heading mw-heading2"><h2 id="Inference_complexity_and_approximation_algorithms">Inference complexity and approximation algorithms</h2></div>
<p>In 1990, while working at Stanford University on large bioinformatic applications, Cooper proved that exact inference in Bayesian networks is <a href="NP-hard" class="mw-redirect" title="NP-hard">NP-hard</a>.<sup id="cite_ref-22" class="reference"><a href="#cite_note-22"><span class="cite-bracket">[</span>22<span class="cite-bracket">]</span></a></sup> This result prompted research on approximation algorithms with the aim of developing a tractable approximation to probabilistic inference. In 1993, Paul Dagum and <a href="Michael_Luby" title="Michael Luby">Michael Luby</a> proved two surprising results on the complexity of approximation of probabilistic inference in Bayesian networks.<sup id="cite_ref-23" class="reference"><a href="#cite_note-23"><span class="cite-bracket">[</span>23<span class="cite-bracket">]</span></a></sup> First, they proved that no tractable <a href="Deterministic_algorithm" title="Deterministic algorithm">deterministic algorithm</a> can approximate probabilistic inference to within an <a href="Absolute_error" class="mw-redirect" title="Absolute error">absolute error</a> <i>ɛ</i> < 1/2. Second, they proved that no tractable <a href="Randomized_algorithm" title="Randomized algorithm">randomized algorithm</a> can approximate probabilistic inference to within an absolute error <i>ɛ</i> < 1/2 with confidence probability greater than 1/2.
</p><p>At about the same time, <a href="Dan_Roth" title="Dan Roth">Roth</a> proved that exact inference in Bayesian networks is in fact <a href="Sharp-P-complete" class="mw-redirect" title="Sharp-P-complete">#P-complete</a> (and thus as hard as counting the number of satisfying assignments of a <a href="Conjunctive_normal_form" title="Conjunctive normal form">conjunctive normal form</a> formula (CNF)) and that approximate inference within a factor 2<sup><i>n</i><sup>1−<i>ɛ</i></sup></sup> for every <i>ɛ</i> > 0, even for Bayesian networks with restricted architecture, is NP-hard.<sup id="cite_ref-24" class="reference"><a href="#cite_note-24"><span class="cite-bracket">[</span>24<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-25" class="reference"><a href="#cite_note-25"><span class="cite-bracket">[</span>25<span class="cite-bracket">]</span></a></sup>
</p><p>In practical terms, these complexity results suggested that while Bayesian networks were rich representations for AI and machine learning applications, their use in large real-world applications would need to be tempered by either topological structural constraints, such as naïve Bayes networks, or by restrictions on the conditional probabilities. The bounded variance algorithm<sup id="cite_ref-26" class="reference"><a href="#cite_note-26"><span class="cite-bracket">[</span>26<span class="cite-bracket">]</span></a></sup> developed by Dagum and Luby was the first provable fast approximation algorithm to efficiently approximate probabilistic inference in Bayesian networks with guarantees on the error approximation. This powerful algorithm required the minor restriction on the conditional probabilities of the Bayesian network to be bounded away from zero and one by <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle 1/p(n)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mn>1</mn>
<mrow class="MJX-TeXAtom-ORD">
<mo>/</mo>
</mrow>
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>n</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle 1/p(n)}</annotation>
</semantics>
</math></span><img src="./97e7cbca3ba3485ca77e70558b757f0a0d79719c.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:6.698ex; height:2.843ex;" alt="{\displaystyle 1/p(n)}" loading="lazy"></span> where <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p(n)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>p</mi>
<mo stretchy="false">(</mo>
<mi>n</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p(n)}</annotation>
</semantics>
</math></span><img src="./0e7d5ae3aa9524f57fb8b44ac46ee8cf6a52d7e6.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; margin-left: -0.089ex; width:4.463ex; height:2.843ex;" alt="{\displaystyle p(n)}" loading="lazy"></span> was any polynomial of the number of nodes in the network, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle n}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>n</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle n}</annotation>
</semantics>
</math></span><img src="./a601995d55609f2d9f5e233e36fbe9ea26011b3b.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.395ex; height:1.676ex;" alt="{\displaystyle n}" loading="lazy"></span>.
</p>
<div class="mw-heading mw-heading2"><h2 id="Software">Software</h2></div>
<p>Notable software for Bayesian networks include:
</p>
<ul><li><a href="Just_another_Gibbs_sampler" title="Just another Gibbs sampler">Just another Gibbs sampler</a> (JAGS) – Open-source alternative to WinBUGS. Uses Gibbs sampling.</li>
<li><a href="OpenBUGS" title="OpenBUGS">OpenBUGS</a> – Open-source development of WinBUGS.</li>
<li><a href="SPSS_Modeler" title="SPSS Modeler">SPSS Modeler</a> – Commercial software that includes an implementation for Bayesian networks.</li>
<li><a href="Stan_(software)" title="Stan (software)">Stan (software)</a> – Stan is an open-source package for obtaining Bayesian inference using the No-U-Turn sampler (NUTS),<sup id="cite_ref-27" class="reference"><a href="#cite_note-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup> a variant of Hamiltonian Monte Carlo.</li>
<li><a href="PyMC" title="PyMC">PyMC</a> – A Python library implementing an embedded domain specific language to represent bayesian networks, and a variety of samplers (including NUTS)</li>
<li><a href="WinBUGS" title="WinBUGS">WinBUGS</a> – One of the first computational implementations of MCMC samplers. No longer maintained.</li></ul>
<div class="mw-heading mw-heading2"><h2 id="History">History</h2></div>
<p>The term Bayesian network was coined by <a href="Judea_Pearl" title="Judea Pearl">Judea Pearl</a> in 1985 to emphasize:<sup id="cite_ref-28" class="reference"><a href="#cite_note-28"><span class="cite-bracket">[</span>28<span class="cite-bracket">]</span></a></sup>
</p>
<ul><li>the often subjective nature of the input information</li>
<li>the reliance on Bayes' conditioning as the basis for updating information</li>
<li>the distinction between causal and evidential modes of reasoning<sup id="cite_ref-29" class="reference"><a href="#cite_note-29"><span class="cite-bracket">[</span>29<span class="cite-bracket">]</span></a></sup></li></ul>
<p>In the late 1980s Pearl's <i>Probabilistic Reasoning in Intelligent Systems</i><sup id="cite_ref-30" class="reference"><a href="#cite_note-30"><span class="cite-bracket">[</span>30<span class="cite-bracket">]</span></a></sup> and <a href="Richard_E._Neapolitan" class="mw-redirect" title="Richard E. Neapolitan">Neapolitan</a>'s <i>Probabilistic Reasoning in Expert Systems</i><sup id="cite_ref-31" class="reference"><a href="#cite_note-31"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup> summarized their properties and established them as a field of study.
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1266661725">
/* start https://en.wikipedia.org/ */
.mw-parser-output .portalbox{padding:0;margin:0.5em 0;display:table;box-sizing:border-box;max-width:175px;list-style:none}.mw-parser-output .portalborder{border:1px solid var(--border-color-base,#a2a9b1);padding:0.1em;background:var(--background-color-neutral-subtle,#f8f9fa)}.mw-parser-output .portalbox-entry{display:table-row;font-size:85%;line-height:110%;height:1.9em;font-style:italic;font-weight:bold}.mw-parser-output .portalbox-image{display:table-cell;padding:0.2em;vertical-align:middle;text-align:center}.mw-parser-output .portalbox-link{display:table-cell;padding:0.2em 0.2em 0.2em 0.3em;vertical-align:middle}@media(min-width:720px){.mw-parser-output .portalleft{margin:0.5em 1em 0.5em 0}.mw-parser-output .portalright{clear:right;float:right;margin:0.5em 0 0.5em 1em}}
/* end https://en.wikipedia.org/ */
</style>
<style data-mw-deduplicate="TemplateStyles:r1184024115">
/* start https://en.wikipedia.org/ */
.mw-parser-output .div-col{margin-top:0.3em;column-width:30em}.mw-parser-output .div-col-small{font-size:90%}.mw-parser-output .div-col-rules{column-rule:1px solid #aaa}.mw-parser-output .div-col dl,.mw-parser-output .div-col ol,.mw-parser-output .div-col ul{margin-top:0}.mw-parser-output .div-col li,.mw-parser-output .div-col dd{page-break-inside:avoid;break-inside:avoid-column}
/* end https://en.wikipedia.org/ */
</style><div class="div-col" style="column-width: 30em;">
<ul><li><a href="Bayesian_epistemology" title="Bayesian epistemology">Bayesian epistemology</a></li>
<li><a href="Bayesian_programming" title="Bayesian programming">Bayesian programming</a></li>
<li><a href="Causal_inference" title="Causal inference">Causal inference</a></li>
<li><a href="Causal_loop_diagram" title="Causal loop diagram">Causal loop diagram</a></li>
<li><a href="Chow%E2%80%93Liu_tree" title="Chow–Liu tree">Chow–Liu tree</a></li>
<li><a href="Computational_intelligence" title="Computational intelligence">Computational intelligence</a></li>
<li><a href="Computational_phylogenetics" title="Computational phylogenetics">Computational phylogenetics</a></li>
<li><a href="Deep_belief_network" title="Deep belief network">Deep belief network</a></li>
<li><a href="Dempster%E2%80%93Shafer_theory" title="Dempster–Shafer theory">Dempster–Shafer theory</a> – a generalization of Bayes' theorem</li>
<li><a href="Expectation%E2%80%93maximization_algorithm" title="Expectation–maximization algorithm">Expectation–maximization algorithm</a></li>
<li><a href="Factor_graph" title="Factor graph">Factor graph</a></li>
<li><a href="Hierarchical_temporal_memory" title="Hierarchical temporal memory">Hierarchical temporal memory</a></li>
<li><a href="Kalman_filter" title="Kalman filter">Kalman filter</a></li>
<li><a href="Memory-prediction_framework" title="Memory-prediction framework">Memory-prediction framework</a></li>
<li><a href="Mixture_distribution" title="Mixture distribution">Mixture distribution</a></li>
<li><a href="Mixture_model" title="Mixture model">Mixture model</a></li>
<li><a href="Naive_Bayes_classifier" title="Naive Bayes classifier">Naive Bayes classifier</a></li>
<li><a href="Plate_notation" title="Plate notation">Plate notation</a></li>
<li><a href="Polytree" title="Polytree">Polytree</a></li>
<li><a href="Sensor_fusion" title="Sensor fusion">Sensor fusion</a></li>
<li><a href="Sequence_alignment" title="Sequence alignment">Sequence alignment</a></li>
<li><a href="Structural_equation_modeling" title="Structural equation modeling">Structural equation modeling</a></li>
<li><a href="Subjective_logic" title="Subjective logic">Subjective logic</a></li>
<li><a href="Variable-order_Bayesian_network" title="Variable-order Bayesian network">Variable-order Bayesian network</a></li></ul></div>
<div class="mw-heading mw-heading2"><h2 id="Notes">Notes</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */
.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}
/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap mw-references-columns"><ol class="references">
<li id="cite_note-1"><span class="mw-cite-backlink"><b><a href="#cite_ref-1">^</a></b></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */
.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}
/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFRuggeriKenettFaltin2007" class="citation book cs1">Ruggeri, Fabrizio; Kenett, Ron S.; Faltin, Frederick W., eds. (2007-12-14). <a rel="nofollow" class="external text" href="https://onlinelibrary.wiley.com/doi/book/10.1002/9780470061572"><i>Encyclopedia of Statistics in Quality and Reliability</i></a> (1 ed.). Wiley. p. 1. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1002%2F9780470061572.eqr089">10.1002/9780470061572.eqr089</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-470-01861-3</bdi>.</cite></span>
</li>
<li id="cite_note-pearl2000-2"><span class="mw-cite-backlink">^ <a href="#cite_ref-pearl2000_2-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-pearl2000_2-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-pearl2000_2-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-pearl2000_2-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-pearl2000_2-4"><sup><i><b>e</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFPearl2000" class="citation book cs1"><a href="Judea_Pearl" title="Judea Pearl">Pearl, Judea</a> (2000). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=LLkhAwAAQBAJ"><i>Causality: Models, Reasoning, and Inference</i></a>. <a href="Cambridge_University_Press" title="Cambridge University Press">Cambridge University Press</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-521-77362-1</bdi>. <a href="OCLC_(identifier)" class="mw-redirect" title="OCLC (identifier)">OCLC</a> <a rel="nofollow" class="external text" href="https://search.worldcat.org/oclc/42291253">42291253</a>.</cite></span>
</li>
<li id="cite_note-3"><span class="mw-cite-backlink"><b><a href="#cite_ref-3">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="http://bayes.cs.ucla.edu/BOOK-2K/ch3-3.pdf">"The Back-Door Criterion"</a> <span class="cs1-format">(PDF)</span><span class="reference-accessdate">. Retrieved <span class="nowrap">2014-09-18</span></span>.</cite></span>
</li>
<li id="cite_note-4"><span class="mw-cite-backlink"><b><a href="#cite_ref-4">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="http://bayes.cs.ucla.edu/BOOK-09/ch11-1-2-final.pdf">"d-Separation without Tears"</a> <span class="cs1-format">(PDF)</span><span class="reference-accessdate">. Retrieved <span class="nowrap">2014-09-18</span></span>.</cite></span>
</li>
<li id="cite_note-pearl-r212-5"><span class="mw-cite-backlink"><b><a href="#cite_ref-pearl-r212_5-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFPearl1994" class="citation conference cs1">Pearl J (1994). <a rel="nofollow" class="external text" href="http://dl.acm.org/ft_gateway.cfm?id=2074452&ftid=1062250&dwn=1&CFID=161588115&CFTOKEN=10243006">"A Probabilistic Calculus of Actions"</a>. In Lopez de Mantaras R, Poole D (eds.). <i>UAI'94 Proceedings of the Tenth international conference on Uncertainty in artificial intelligence</i>. San Mateo CA: <a href="Morgan_Kaufmann" class="mw-redirect" title="Morgan Kaufmann">Morgan Kaufmann</a>. pp. <span class="nowrap">454–</span>462. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1302.6835">1302.6835</a></span>. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2013arXiv1302.6835P">2013arXiv1302.6835P</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>1-55860-332-8</bdi>.</cite></span>
</li>
<li id="cite_note-6"><span class="mw-cite-backlink"><b><a href="#cite_ref-6">^</a></b></span> <span class="reference-text"><cite id="CITEREFShpitserPearl2006" class="citation book cs1">Shpitser I, Pearl J (2006). "Identification of Conditional Interventional Distributions". In Dechter R, Richardson TS (eds.). <i>Proceedings of the Twenty-Second Conference on Uncertainty in Artificial Intelligence</i>. Corvallis, OR: AUAI Press. pp. <span class="nowrap">437–</span>444. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1206.6876">1206.6876</a></span>.</cite></span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><b><a href="#cite_ref-7">^</a></b></span> <span class="reference-text"><cite id="CITEREFRebanePearl1987" class="citation book cs1">Rebane G, Pearl J (1987). "The Recovery of Causal Poly-trees from Statistical Data". <i>Proceedings, 3rd Workshop on Uncertainty in AI</i>. Seattle, WA. pp. <span class="nowrap">222–</span>228. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1304.2736">1304.2736</a></span>.</cite><span class="cs1-maint citation-comment"><code class="cs1-code">{{cite book}}</code>: CS1 maint: location missing publisher (link)</span></span>
</li>
<li id="cite_note-8"><span class="mw-cite-backlink"><b><a href="#cite_ref-8">^</a></b></span> <span class="reference-text"><cite id="CITEREFSpirtesGlymour1991" class="citation journal cs1">Spirtes P, Glymour C (1991). <a rel="nofollow" class="external text" href="http://repository.cmu.edu/cgi/viewcontent.cgi?article=1316&context=philosophy">"An algorithm for fast recovery of sparse causal graphs"</a> <span class="cs1-format">(PDF)</span>. <i>Social Science Computer Review</i>. <b>9</b> (1): <span class="nowrap">62–</span>72. <a href="CiteSeerX_(identifier)" class="mw-redirect" title="CiteSeerX (identifier)">CiteSeerX</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.650.2922">10.1.1.650.2922</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1177%2F089443939100900106">10.1177/089443939100900106</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a> <a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:38398322">38398322</a>.</cite></span>
</li>
<li id="cite_note-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-9">^</a></b></span> <span class="reference-text"><cite id="CITEREFSpirtesGlymourScheines1993" class="citation book cs1">Spirtes P, Glymour CN, Scheines R (1993). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=VkawQgAACAAJ"><i>Causation, Prediction, and Search</i></a> (1st ed.). Springer-Verlag. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-387-97979-3</bdi>.</cite></span>
</li>
<li id="cite_note-10"><span class="mw-cite-backlink"><b><a href="#cite_ref-10">^</a></b></span> <span class="reference-text"><cite id="CITEREFVermaPearl1991" class="citation conference cs1">Verma T, Pearl J (1991). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=ikuuHAAACAAJ">"Equivalence and synthesis of causal models"</a>. In Bonissone P, Henrion M, Kanal LN, Lemmer JF (eds.). <i>UAI '90 Proceedings of the Sixth Annual Conference on Uncertainty in Artificial Intelligence</i>. Elsevier. pp. <span class="nowrap">255–</span>270. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>0-444-89264-8</bdi>.</cite></span>
</li>
<li id="cite_note-11"><span class="mw-cite-backlink"><b><a href="#cite_ref-11">^</a></b></span> <span class="reference-text"><cite id="CITEREFFriedmanGeigerGoldszmidt1997" class="citation journal cs1">Friedman N, Geiger D, Goldszmidt M (November 1997). <a rel="nofollow" class="external text" href="https://doi.org/10.1023%2FA%3A1007465528199">"Bayesian Network Classifiers"</a>. <i>Machine Learning</i>. <b>29</b> (<span class="nowrap">2–</span>3): <span class="nowrap">131–</span>163. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1023%2FA%3A1007465528199">10.1023/A:1007465528199</a></span>.</cite></span>
</li>
<li id="cite_note-12"><span class="mw-cite-backlink"><b><a href="#cite_ref-12">^</a></b></span> <span class="reference-text"><cite id="CITEREFFriedmanLinialNachmanPe'er2000" class="citation journal cs1">Friedman N, Linial M, Nachman I, Pe'er D (August 2000). "Using Bayesian networks to analyze expression data". <i>Journal of Computational Biology</i>. <b>7</b> (<span class="nowrap">3–</span>4): <span class="nowrap">601–</span>20. <a href="CiteSeerX_(identifier)" class="mw-redirect" title="CiteSeerX (identifier)">CiteSeerX</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.191.139">10.1.1.191.139</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1089%2F106652700750050961">10.1089/106652700750050961</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/11108481">11108481</a>.</cite></span>
</li>
<li id="cite_note-13"><span class="mw-cite-backlink"><b><a href="#cite_ref-13">^</a></b></span> <span class="reference-text"><cite id="CITEREFCussens2011" class="citation journal cs1 cs1-prop-unfit">Cussens J (2011). <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220327163338/https://dslpitt.org/papers/11/p153-cussens.pdf">"Bayesian network learning with cutting planes"</a> <span class="cs1-format">(PDF)</span>. <i>Proceedings of the 27th Conference Annual Conference on Uncertainty in Artificial Intelligence</i>: <span class="nowrap">153–</span>160. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1202.3713">1202.3713</a></span>. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2012arXiv1202.3713C">2012arXiv1202.3713C</a>. Archived from the original on March 27, 2022.</cite></span>
</li>
<li id="cite_note-14"><span class="mw-cite-backlink"><b><a href="#cite_ref-14">^</a></b></span> <span class="reference-text"><cite id="CITEREFScanagattade_CamposCoraniZaffalon2015" class="citation book cs1">Scanagatta M, de Campos CP, Corani G, Zaffalon M (2015). <a rel="nofollow" class="external text" href="https://papers.nips.cc/paper/5803-learning-bayesian-networks-with-thousands-of-variables">"Learning Bayesian Networks with Thousands of Variables"</a>. <i>NIPS-15: Advances in Neural Information Processing Systems</i>. Vol. 28. Curran Associates. pp. <span class="nowrap">1855–</span>1863.</cite></span>
</li>
<li id="cite_note-Petitjean-15"><span class="mw-cite-backlink"><b><a href="#cite_ref-Petitjean_15-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFPetitjeanWebbNicholson2013" class="citation conference cs1">Petitjean F, Webb GI, Nicholson AE (2013). <a rel="nofollow" class="external text" href="http://www.tiny-clues.eu/Research/Petitjean2013-ICDM.pdf"><i>Scaling log-linear analysis to high-dimensional data</i></a> <span class="cs1-format">(PDF)</span>. International Conference on Data Mining. Dallas, TX, USA: IEEE.</cite></span>
</li>
<li id="cite_note-16"><span class="mw-cite-backlink"><b><a href="#cite_ref-16">^</a></b></span> <span class="reference-text">M. Scanagatta, G. Corani, C. P. de Campos, and M. Zaffalon. <a rel="nofollow" class="external text" href="http://papers.nips.cc/paper/6232-learning-treewidth-bounded-bayesian-networks-with-thousands-of-variables">Learning Treewidth-Bounded Bayesian Networks with Thousands of Variables.</a> In NIPS-16: Advances in Neural Information Processing Systems 29, 2016.</span>
</li>
<li id="cite_note-FOOTNOTERussellNorvig2003496-17"><span class="mw-cite-backlink">^ <a href="#cite_ref-FOOTNOTERussellNorvig2003496_17-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-FOOTNOTERussellNorvig2003496_17-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><a href="#CITEREFRussellNorvig2003">Russell & Norvig 2003</a>, p. 496.</span>
</li>
<li id="cite_note-FOOTNOTERussellNorvig2003499-18"><span class="mw-cite-backlink">^ <a href="#cite_ref-FOOTNOTERussellNorvig2003499_18-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-FOOTNOTERussellNorvig2003499_18-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><a href="#CITEREFRussellNorvig2003">Russell & Norvig 2003</a>, p. 499.</span>
</li>
<li id="cite_note-19"><span class="mw-cite-backlink"><b><a href="#cite_ref-19">^</a></b></span> <span class="reference-text"><cite id="CITEREFChickeringHeckermanMeek2004" class="citation journal cs1">Chickering, David M.; Heckerman, David; Meek, Christopher (2004). <a rel="nofollow" class="external text" href="https://www.jmlr.org/papers/volume5/chickering04a/chickering04a.pdf">"Large-sample learning of Bayesian networks is NP-hard"</a> <span class="cs1-format">(PDF)</span>. <i>Journal of Machine Learning Research</i>. <b>5</b>: <span class="nowrap">1287–</span>1330.</cite></span>
</li>
<li id="cite_note-20"><span class="mw-cite-backlink"><b><a href="#cite_ref-20">^</a></b></span> <span class="reference-text"><cite id="CITEREFDeligeorgakiMarkhamMisraSolus2023" class="citation journal cs1">Deligeorgaki, Danai; Markham, Alex; Misra, Pratik; Solus, Liam (2023). "Combinatorial and algebraic perspectives on the marginal independence structure of Bayesian networks". <i>Algebraic Statistics</i>. <b>14</b> (2): <span class="nowrap">233–</span>286. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2210.00822">2210.00822</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.2140%2Fastat.2023.14.233">10.2140/astat.2023.14.233</a>.</cite></span>
</li>
<li id="cite_note-21"><span class="mw-cite-backlink"><b><a href="#cite_ref-21">^</a></b></span> <span class="reference-text"><cite id="CITEREFNeapolitan2004" class="citation book cs1">Neapolitan RE (2004). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=OlMZAQAAIAAJ"><i>Learning Bayesian networks</i></a>. Prentice Hall. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-13-012534-7</bdi>.</cite></span>
</li>
<li id="cite_note-22"><span class="mw-cite-backlink"><b><a href="#cite_ref-22">^</a></b></span> <span class="reference-text">
<cite id="CITEREFCooper1990" class="citation journal cs1">Cooper GF (1990). <a rel="nofollow" class="external text" href="https://stat.duke.edu/~sayan/npcomplete.pdf">"The Computational Complexity of Probabilistic Inference Using Bayesian Belief Networks"</a> <span class="cs1-format">(PDF)</span>. <i>Artificial Intelligence</i>. <b>42</b> (<span class="nowrap">2–</span>3): <span class="nowrap">393–</span>405. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2F0004-3702%2890%2990060-d">10.1016/0004-3702(90)90060-d</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a> <a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:43363498">43363498</a>.</cite></span>
</li>
<li id="cite_note-23"><span class="mw-cite-backlink"><b><a href="#cite_ref-23">^</a></b></span> <span class="reference-text">
<cite id="CITEREFDagumLuby1993" class="citation journal cs1">Dagum P, <a href="Michael_Luby" title="Michael Luby">Luby M</a> (1993). "Approximating probabilistic inference in Bayesian belief networks is NP-hard". <i>Artificial Intelligence</i>. <b>60</b> (1): <span class="nowrap">141–</span>153. <a href="CiteSeerX_(identifier)" class="mw-redirect" title="CiteSeerX (identifier)">CiteSeerX</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.333.1586">10.1.1.333.1586</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2F0004-3702%2893%2990036-b">10.1016/0004-3702(93)90036-b</a>.</cite></span>
</li>
<li id="cite_note-24"><span class="mw-cite-backlink"><b><a href="#cite_ref-24">^</a></b></span> <span class="reference-text">D. Roth, <a rel="nofollow" class="external text" href="http://cogcomp.cs.illinois.edu/page/publication_view/5">On the hardness of approximate reasoning</a>, IJCAI (1993)</span>
</li>
<li id="cite_note-25"><span class="mw-cite-backlink"><b><a href="#cite_ref-25">^</a></b></span> <span class="reference-text">D. Roth, <a rel="nofollow" class="external text" href="http://cogcomp.cs.illinois.edu/papers/hardJ.pdf">On the hardness of approximate reasoning</a>, Artificial Intelligence (1996)</span>
</li>
<li id="cite_note-26"><span class="mw-cite-backlink"><b><a href="#cite_ref-26">^</a></b></span> <span class="reference-text"><cite id="CITEREFDagumLuby1997" class="citation journal cs1">Dagum P, <a href="Michael_Luby" title="Michael Luby">Luby M</a> (1997). <a rel="nofollow" class="external text" href="https://web.archive.org/web/20170706064354/http://www1.icsi.berkeley.edu/~luby/PAPERS/bayesian.ps">"An optimal approximation algorithm for Bayesian inference"</a>. <i>Artificial Intelligence</i>. <b>93</b> (<span class="nowrap">1–</span>2): <span class="nowrap">1–</span>27. <a href="CiteSeerX_(identifier)" class="mw-redirect" title="CiteSeerX (identifier)">CiteSeerX</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.36.7946">10.1.1.36.7946</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fs0004-3702%2897%2900013-1">10.1016/s0004-3702(97)00013-1</a>. Archived from <a rel="nofollow" class="external text" href="http://icsi.berkeley.edu/~luby/PAPERS/bayesian.ps">the original</a> on 2017-07-06<span class="reference-accessdate">. Retrieved <span class="nowrap">2015-12-19</span></span>.</cite></span>
</li>
<li id="cite_note-27"><span class="mw-cite-backlink"><b><a href="#cite_ref-27">^</a></b></span> <span class="reference-text"><cite id="CITEREFHoffmanGelman2011" class="citation arxiv cs1">Hoffman, Matthew D.; Gelman, Andrew (2011). "The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1111.4246">1111.4246</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/stat.CO">stat.CO</a>].</cite></span>
</li>
<li id="cite_note-28"><span class="mw-cite-backlink"><b><a href="#cite_ref-28">^</a></b></span> <span class="reference-text"><cite id="CITEREFPearl1985" class="citation conference cs1"><a href="Judea_Pearl" title="Judea Pearl">Pearl J</a> (1985). <a rel="nofollow" class="external text" href="http://ftp.cs.ucla.edu/tech-report/198_-reports/850017.pdf"><i>Bayesian Networks: A Model of Self-Activated Memory for Evidential Reasoning</i></a> <span class="cs1-format">(UCLA Technical Report CSD-850017)</span>. Proceedings of the 7th Conference of the Cognitive Science Society, University of California, Irvine, CA. pp. <span class="nowrap">329–</span>334<span class="reference-accessdate">. Retrieved <span class="nowrap">2009-05-01</span></span>.</cite></span>
</li>
<li id="cite_note-29"><span class="mw-cite-backlink"><b><a href="#cite_ref-29">^</a></b></span> <span class="reference-text"><cite id="CITEREFBayesPrice1763" class="citation journal cs1"><a href="Thomas_Bayes" title="Thomas Bayes">Bayes T</a>, Price (1763). <a href="An_Essay_Towards_Solving_a_Problem_in_the_Doctrine_of_Chances" title="An Essay Towards Solving a Problem in the Doctrine of Chances">"An Essay Towards Solving a Problem in the Doctrine of Chances"</a>. <i><a href="Philosophical_Transactions_of_the_Royal_Society" title="Philosophical Transactions of the Royal Society">Philosophical Transactions of the Royal Society</a></i>. <b>53</b>: <span class="nowrap">370–</span>418. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1098%2Frstl.1763.0053">10.1098/rstl.1763.0053</a></span>.</cite></span>
</li>
<li id="cite_note-30"><span class="mw-cite-backlink"><b><a href="#cite_ref-30">^</a></b></span> <span class="reference-text"><cite id="CITEREFPearl1988" class="citation book cs1">Pearl J (1988-09-15). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=AvNID7LyMusC"><i>Probabilistic Reasoning in Intelligent Systems</i></a>. San Francisco CA: <a href="Morgan_Kaufmann" class="mw-redirect" title="Morgan Kaufmann">Morgan Kaufmann</a>. p. 1988. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-1-55860-479-7</bdi>.</cite></span>
</li>
<li id="cite_note-31"><span class="mw-cite-backlink"><b><a href="#cite_ref-31">^</a></b></span> <span class="reference-text"><cite id="CITEREFNeapolitan1989" class="citation book cs1">Neapolitan RE (1989). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=7X5KLwEACAAJ"><i>Probabilistic reasoning in expert systems: theory and algorithms</i></a>. Wiley. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-471-61840-9</bdi>.</cite></span>
</li>
</ol></div></div>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239549316">
/* start https://en.wikipedia.org/ */
.mw-parser-output .refbegin{margin-bottom:0.5em}.mw-parser-output .refbegin-hanging-indents>ul{margin-left:0}.mw-parser-output .refbegin-hanging-indents>ul>li{margin-left:0;padding-left:3.2em;text-indent:-3.2em}.mw-parser-output .refbegin-hanging-indents ul,.mw-parser-output .refbegin-hanging-indents ul li{list-style:none}@media(max-width:720px){.mw-parser-output .refbegin-hanging-indents>ul>li{padding-left:1.6em;text-indent:-1.6em}}.mw-parser-output .refbegin-columns{margin-top:0.3em}.mw-parser-output .refbegin-columns ul{margin-top:0}.mw-parser-output .refbegin-columns li{page-break-inside:avoid;break-inside:avoid-column}@media screen{.mw-parser-output .refbegin{font-size:90%}}
/* end https://en.wikipedia.org/ */
</style><div class="refbegin" style="">
<ul><li><cite id="CITEREFBen_Gal2007" class="citation encyclopaedia cs1">Ben Gal I (2007). <a rel="nofollow" class="external text" href="https://web.archive.org/web/20161123041500/http://www.eng.tau.ac.il/~bengal/BN.pdf">"Bayesian Networks"</a> <span class="cs1-format">(PDF)</span>. In Ruggeri F, Kennett RS, Faltin FW (eds.). <i>Support-Page</i>. <i>Encyclopedia of Statistics in Quality and Reliability</i>. <a href="John_Wiley_%26_Sons" class="mw-redirect" title="John Wiley & Sons">John Wiley & Sons</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1002%2F9780470061572.eqr089">10.1002/9780470061572.eqr089</a></span>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-470-01861-3</bdi>. Archived from <a rel="nofollow" class="external text" href="http://www.eng.tau.ac.il/~bengal/BN.pdf">the original</a> <span class="cs1-format">(PDF)</span> on 2016-11-23<span class="reference-accessdate">. Retrieved <span class="nowrap">2007-08-27</span></span>.</cite></li>
<li><cite id="CITEREFBertsch_McGrayne2011" class="citation book cs1">Bertsch McGrayne S (2011). <span class="id-lock-registration" title="Free registration required"><a rel="nofollow" class="external text" href="https://archive.org/details/theorythatwouldn0000mcgr"><i>The Theory That Would not Die</i></a></span>. New Haven: <a href="Yale_University_Press" title="Yale University Press">Yale University Press</a>.</cite></li>
<li><cite id="CITEREFBorgeltKruse2002" class="citation book cs1">Borgelt C, <a href="Rudolf_Kruse" title="Rudolf Kruse">Kruse R</a> (March 2002). <a rel="nofollow" class="external text" href="http://fuzzy.cs.uni-magdeburg.de/books/gm/"><i>Graphical Models: Methods for Data Analysis and Mining</i></a>. <a href="Chichester" title="Chichester">Chichester, UK</a>: <a href="John_Wiley_%26_Sons" class="mw-redirect" title="John Wiley & Sons">Wiley</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-470-84337-6</bdi>.</cite></li>
<li><cite id="CITEREFBorsuk2008" class="citation encyclopaedia cs1">Borsuk ME (2008). "Ecological informatics: Bayesian networks". In <a href="Sven_Erik_J%C3%B8rgensen" title="Sven Erik Jørgensen">Jørgensen, Sven Erik</a>, Fath, Brian (eds.). <i>Encyclopedia of Ecology</i>. Elsevier. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-444-52033-3</bdi>.</cite></li>
<li><cite id="CITEREFCastilloGutiérrezHadi1997" class="citation book cs1">Castillo E, Gutiérrez JM, Hadi AS (1997). "Learning Bayesian Networks". <i>Expert Systems and Probabilistic Network Models</i>. Monographs in computer science. New York: <a href="Springer_Science%2BBusiness_Media" title="Springer Science+Business Media">Springer-Verlag</a>. pp. <span class="nowrap">481–</span>528. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-387-94858-4</bdi>.</cite></li>
<li><cite id="CITEREFComleyDowe2003" class="citation journal cs1">Comley JW, Dowe DL (June 2003). <a rel="nofollow" class="external text" href="http://www.csse.monash.edu.au/~dld/David.Dowe.publications.html#ComleyDowe2003">"General Bayesian networks and asymmetric languages"</a>. <i>Proceedings of the 2nd Hawaii International Conference on Statistics and Related Fields</i>.</cite></li>
<li><cite id="CITEREFComleyDowe2005" class="citation book cs1">Comley JW, Dowe DL (2005). <a rel="nofollow" class="external text" href="http://www.csse.monash.edu.au/~dld/David.Dowe.publications.html#ComleyDowe2005">"Minimum Message Length and Generalized Bayesian Nets with Asymmetric Languages"</a>. In Grünwald PD, Myung IJ, Pitt MA (eds.). <i>Advances in Minimum Description Length: Theory and Applications</i>. Neural information processing series. <a href="Cambridge%2C_Massachusetts" title="Cambridge, Massachusetts">Cambridge, Massachusetts</a>: Bradford Books (<a href="MIT_Press" title="MIT Press">MIT Press</a>) (published April 2005). pp. <span class="nowrap">265–</span>294. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-262-07262-5</bdi>.</cite> (This paper puts <a href="Decision_tree_learning" title="Decision tree learning">decision trees</a> in internal nodes of Bayes networks using <a rel="nofollow" class="external text" href="http://www.csse.monash.edu.au/~dld/MML.html">Minimum Message Length</a> (<a href="Minimum_message_length" title="Minimum message length">MML</a>).</li>
<li><cite id="CITEREFDarwiche2009" class="citation book cs1">Darwiche A (2009). <a rel="nofollow" class="external text" href="http://www.cambridge.org/9780521884389"><i>Modeling and Reasoning with Bayesian Networks</i></a>. <a href="Cambridge_University_Press" title="Cambridge University Press">Cambridge University Press</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-521-88438-9</bdi>.</cite></li>
<li><cite id="CITEREFDowe2011" class="citation book cs1">Dowe, David L. (2011-05-31). <a rel="nofollow" class="external text" href="http://www.csse.monash.edu.au/~dld/Publications/2010/Dowe2010_MML_HandbookPhilSci_Vol7_HandbookPhilStat_MML+hybridBayesianNetworkGraphicalModels+StatisticalConsistency+InvarianceAndUniqueness_pp901-982.pdf">"Hybrid Bayesian network graphical models, statistical consistency, invariance and uniqueness"</a> <span class="cs1-format">(PDF)</span>. <a rel="nofollow" class="external text" href="https://books.google.com/books?id=mPG5RupkTX0C"><i>Philosophy of Statistics</i></a>. Elsevier. pp. <a rel="nofollow" class="external text" href="http://www.csse.monash.edu.au/~dld/Publications/2010/Dowe2010_MML_HandbookPhilSci_Vol7_HandbookPhilStat_MML+hybridBayesianNetworkGraphicalModels+StatisticalConsistency+InvarianceAndUniqueness_pp901-982.pdf">901–982</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-08-093096-1</bdi>.</cite></li>
<li><cite id="CITEREFFentonNeil2007" class="citation book cs1">Fenton N, Neil ME (November 2007). <a rel="nofollow" class="external text" href="https://web.archive.org/web/20080514044436/http://www.agenarisk.com/resources/apps_bayesian_networks.pdf">"Managing Risk in the Modern World: Applications of Bayesian Networks"</a> <span class="cs1-format">(PDF)</span>. <i>A Knowledge Transfer Report from the London Mathematical Society and the Knowledge Transfer Network for Industrial Mathematics</i>. <a href="London" title="London">London (England)</a>: <a href="London_Mathematical_Society" title="London Mathematical Society">London Mathematical Society</a>. Archived from <a rel="nofollow" class="external text" href="http://www.agenarisk.com/resources/apps_bayesian_networks.pdf">the original</a> <span class="cs1-format">(PDF)</span> on 2008-05-14<span class="reference-accessdate">. Retrieved <span class="nowrap">2008-10-29</span></span>.</cite></li>
<li><cite id="CITEREFFentonNeil2004" class="citation news cs1">Fenton N, Neil ME (July 23, 2004). <a rel="nofollow" class="external text" href="https://web.archive.org/web/20070927153751/https://www.dcs.qmul.ac.uk/~norman/papers/Combining%20evidence%20in%20risk%20analysis%20using%20BNs.pdf">"Combining evidence in risk analysis using Bayesian Networks"</a> <span class="cs1-format">(PDF)</span>. <i>Safety Critical Systems Club Newsletter</i>. Vol. 13, no. 4. <a href="Newcastle_upon_Tyne" title="Newcastle upon Tyne">Newcastle upon Tyne</a>, England. pp. <span class="nowrap">8–</span>13. Archived from <a rel="nofollow" class="external text" href="https://www.dcs.qmul.ac.uk/~norman/papers/Combining%20evidence%20in%20risk%20analysis%20using%20BNs.pdf">the original</a> <span class="cs1-format">(PDF)</span> on 2007-09-27.</cite></li>
<li><cite id="CITEREFGelmanCarlinSternRubin2003" class="citation book cs1">Gelman A, Carlin JB, Stern HS, Rubin DB (2003). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=TNYhnkXQSjAC&pg=PA120">"Part II: Fundamentals of Bayesian Data Analysis: Ch.5 Hierarchical models"</a>. <a rel="nofollow" class="external text" href="https://books.google.com/books?id=TNYhnkXQSjAC"><i>Bayesian Data Analysis</i></a>. <a href="CRC_Press" title="CRC Press">CRC Press</a>. pp. 120–. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-1-58488-388-3</bdi>.</cite></li>
<li><cite id="CITEREFHeckerman1995" class="citation book cs1">Heckerman, David (March 1, 1995). <a rel="nofollow" class="external text" href="https://web.archive.org/web/20060719171558/http://research.microsoft.com/research/pubs/view.aspx?msr_tr_id=MSR-TR-95-06">"Tutorial on Learning with Bayesian Networks"</a>. In Jordan, Michael Irwin (ed.). <i>Learning in Graphical Models</i>. Adaptive Computation and Machine Learning. <a href="Cambridge%2C_Massachusetts" title="Cambridge, Massachusetts">Cambridge, Massachusetts</a>: <a href="MIT_Press" title="MIT Press">MIT Press</a> (published 1998). pp. <span class="nowrap">301–</span>354. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-262-60032-3</bdi>. Archived from the original on July 19, 2006<span class="reference-accessdate">. Retrieved <span class="nowrap">September 15,</span> 2006</span>.</cite><span class="cs1-maint citation-comment"><code class="cs1-code">{{cite book}}</code>: CS1 maint: bot: original URL status unknown (link)</span>:Also appears as <cite id="CITEREFHeckerman1997" class="citation journal cs1">Heckerman, David (March 1997). "Bayesian Networks for Data Mining". <i><a href="Data_Mining_and_Knowledge_Discovery" title="Data Mining and Knowledge Discovery">Data Mining and Knowledge Discovery</a></i>. <b>1</b> (1): <span class="nowrap">79–</span>119. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1023%2FA%3A1009730122752">10.1023/A:1009730122752</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a> <a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:6294315">6294315</a>.</cite></li></ul>
<dl><dd>An earlier version appears as, <a href="Microsoft_Research" title="Microsoft Research">Microsoft Research</a>, March 1, 1995. The paper is about both parameter and structure learning in Bayesian networks.</dd></dl>
<ul><li><cite id="CITEREFJensenNielsen2007" class="citation book cs1">Jensen FV, Nielsen TD (June 6, 2007). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=cWLaBwAAQBAJ"><i>Bayesian Networks and Decision Graphs</i></a>. Information Science and Statistics series (2nd ed.). <a href="New_York_City" title="New York City">New York</a>: <a href="Springer_Science%2BBusiness_Media" title="Springer Science+Business Media">Springer-Verlag</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-387-68281-5</bdi>.</cite></li>
<li><cite id="CITEREFKarimiHamilton2000" class="citation journal cs1">Karimi K, Hamilton HJ (2000). <a rel="nofollow" class="external text" href="http://www.kamran-karimi.com/pubs/khISMIS2000.pdf">"Finding temporal relations: Causal bayesian networks vs. C4. 5"</a> <span class="cs1-format">(PDF)</span>. <i>Twelfth International Symposium on Methodologies for Intelligent Systems</i>.</cite></li>
<li><cite id="CITEREFKorbNicholson2010" class="citation book cs1">Korb KB, Nicholson AE (December 2010). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=LxXOBQAAQBAJ"><i>Bayesian Artificial Intelligence</i></a>. CRC Computer Science & Data Analysis (2nd ed.). <a href="Chapman_%26_Hall" title="Chapman & Hall">Chapman & Hall</a> (<a href="CRC_Press" title="CRC Press">CRC Press</a>). <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs10044-004-0214-5">10.1007/s10044-004-0214-5</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-1-58488-387-6</bdi>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a> <a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:22138783">22138783</a>.</cite></li>
<li><cite id="CITEREFLunnSpiegelhalterThomasBest2009" class="citation journal cs1">Lunn D, Spiegelhalter D, Thomas A, Best N (November 2009). "The BUGS project: Evolution, critique and future directions". <i>Statistics in Medicine</i>. <b>28</b> (25): <span class="nowrap">3049–</span>67. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1002%2Fsim.3680">10.1002/sim.3680</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/19630097">19630097</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a> <a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:7717482">7717482</a>.</cite></li>
<li><cite id="CITEREFNeilFentonTailor2005" class="citation journal cs1">Neil M, Fenton N, Tailor M (August 2005). Greenberg, Michael R. (ed.). <a rel="nofollow" class="external text" href="http://www.dcs.qmul.ac.uk/~norman/papers/oprisk.pdf">"Using Bayesian networks to model expected and unexpected operational losses"</a> <span class="cs1-format">(PDF)</span>. <i>Risk Analysis</i>. <b>25</b> (4): <span class="nowrap">963–</span>72. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2005RiskA..25..963N">2005RiskA..25..963N</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1111%2Fj.1539-6924.2005.00641.x">10.1111/j.1539-6924.2005.00641.x</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/16268944">16268944</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a> <a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:3254505">3254505</a>.</cite></li>
<li><cite id="CITEREFPearl1986" class="citation journal cs1"><a href="Judea_Pearl" title="Judea Pearl">Pearl J</a> (September 1986). "Fusion, propagation, and structuring in belief networks". <i><a href="Artificial_Intelligence_(journal)" title="Artificial Intelligence (journal)">Artificial Intelligence</a></i>. <b>29</b> (3): <span class="nowrap">241–</span>288. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2F0004-3702%2886%2990072-X">10.1016/0004-3702(86)90072-X</a>.</cite></li>
<li><cite id="CITEREFPearl1988" class="citation book cs1"><a href="Judea_Pearl" title="Judea Pearl">Pearl J</a> (1988). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=mn2jBQAAQBAJ"><i>Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference</i></a>. Representation and Reasoning Series (2nd printing ed.). <a href="San_Francisco%2C_California" class="mw-redirect" title="San Francisco, California">San Francisco, California</a>: <a href="Morgan_Kaufmann" class="mw-redirect" title="Morgan Kaufmann">Morgan Kaufmann</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-934613-73-6</bdi>.</cite></li>
<li><cite id="CITEREFPearlRussell2002" class="citation book cs1"><a href="Judea_Pearl" title="Judea Pearl">Pearl J</a>, <a href="Stuart_J._Russell" title="Stuart J. Russell">Russell S</a> (November 2002). "Bayesian Networks". In <a href="Michael_A._Arbib" title="Michael A. Arbib">Arbib MA</a> (ed.). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=Av6qWhtw0-EC"><i>Handbook of Brain Theory and Neural Networks</i></a>. <a href="Cambridge%2C_Massachusetts" title="Cambridge, Massachusetts">Cambridge, Massachusetts</a>: Bradford Books (<a href="MIT_Press" title="MIT Press">MIT Press</a>). pp. <span class="nowrap">157–</span>160. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-262-01197-6</bdi>.</cite></li>
<li><cite id="CITEREFRussellNorvig2003" class="citation cs2"><a href="Stuart_J._Russell" title="Stuart J. Russell">Russell, Stuart J.</a>; <a href="Peter_Norvig" title="Peter Norvig">Norvig, Peter</a> (2003), <a rel="nofollow" class="external text" href="http://aima.cs.berkeley.edu/"><i>Artificial Intelligence: A Modern Approach</i></a> (2nd ed.), Upper Saddle River, New Jersey: Prentice Hall, <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>0-13-790395-2</bdi></cite>.</li>
<li><cite id="CITEREFZhangPoole1994" class="citation journal cs1">Zhang NL, Poole D (May 1994). <a rel="nofollow" class="external text" href="http://www.cs.ust.hk/~lzhang/paper/pspdf/canai94.pdf">"A simple approach to Bayesian network computations"</a> <span class="cs1-format">(PDF)</span>. <i>Proceedings of the Tenth Biennial Canadian Artificial Intelligence Conference (AI-94).</i>: <span class="nowrap">171–</span>178.</cite> This paper presents variable elimination for belief networks.</li></ul>
</div>
<div class="mw-heading mw-heading2"><h2 id="Further_reading">Further reading</h2></div>
<ul><li><cite id="CITEREFConradyJouffe2015" class="citation book cs1">Conrady S, Jouffe L (2015-07-01). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=etXXsgEACAAJ"><i>Bayesian Networks and BayesiaLab – A practical introduction for researchers</i></a>. Franklin, Tennessee: Bayesian USA. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-9965333-0-0</bdi>.</cite></li>
<li><cite id="CITEREFCharniak1991" class="citation web cs1">Charniak E (Winter 1991). <a rel="nofollow" class="external text" href="http://pages.cs.wisc.edu/~dyer/cs540/handouts/charniak.pdf">"Bayesian networks without tears"</a> <span class="cs1-format">(PDF)</span>. <i>AI Magazine</i>.</cite></li>
<li><cite id="CITEREFKruseBorgeltKlawonnMoewes2013" class="citation book cs1">Kruse R, Borgelt C, Klawonn F, Moewes C, Steinbrecher M, Held P (2013). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=etXXsgEACAAJ"><i>Computational Intelligence A Methodological Introduction</i></a>. London: Springer-Verlag. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-1-4471-5012-1</bdi>.</cite></li>
<li><cite id="CITEREFBorgeltSteinbrecherKruse2009" class="citation book cs1">Borgelt C, Steinbrecher M, Kruse R (2009). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=I8Fa-LKDpF0C"><i>Graphical Models – Representations for Learning, Reasoning and Data Mining</i></a> (Second ed.). Chichester: Wiley. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-0-470-74956-2</bdi>.</cite></li></ul>
<div class="mw-heading mw-heading2"><h2 id="External_links">External links</h2></div>
<ul><li><a rel="nofollow" class="external text" href="http://www.niedermayer.ca/papers/bayesian/bayes.html">An Introduction to Bayesian Networks and their Contemporary Applications</a></li>
<li><a rel="nofollow" class="external text" href="http://www.dcs.qmw.ac.uk/%7Enorman/BBNs/BBNs.htm">On-line Tutorial on Bayesian nets and probability</a></li>
<li><a rel="nofollow" class="external text" href="https://web.archive.org/web/20170601002137/http://princesofserendib.com/">Web-App to create Bayesian nets and run it with a Monte Carlo method</a></li>
<li><a rel="nofollow" class="external text" href="http://robotics.stanford.edu/~nodelman/papers/ctbn.pdf">Continuous Time Bayesian Networks</a></li>
<li><a rel="nofollow" class="external text" href="https://web.archive.org/web/20090923200511/http://wiki.syncleus.com/index.php/DANN%3ABayesian_Network">Bayesian Networks: Explanation and Analogy</a></li>
<li><a rel="nofollow" class="external text" href="http://videolectures.net/kdd07_neapolitan_lbn/">A live tutorial on learning Bayesian networks</a></li>
<li><a rel="nofollow" class="external text" href="http://www.biomedcentral.com/1471-2105/7/514/abstract">A hierarchical Bayes Model for handling sample heterogeneity in classification problems</a>, provides a classification model taking into consideration the uncertainty associated with measuring replicate samples.</li>
<li><a rel="nofollow" class="external text" href="http://www.labmedinfo.org/download/lmi339.pdf">Hierarchical Naive Bayes Model for handling sample uncertainty</a> <a rel="nofollow" class="external text" href="https://web.archive.org/web/20070928081740/http://www.labmedinfo.org/download/lmi339.pdf">Archived</a> 2007-09-28 at the <a href="Wayback_Machine" title="Wayback Machine">Wayback Machine</a>, shows how to perform classification and learning with continuous and discrete variables with replicated measurements.</li></ul></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-04-04" href="https://en.wikipedia.org/wiki/?title=Bayesian_network&oldid=1283980484">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>
</body></html>